When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning
Ranking
Overall
80
Content
100
Popularity
34
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper explains why benign fine-tuning can rapidly erode LLM refusal behavior: safety alignment relies on a low-rank output-routing mechanism that is easily disrupted. The findings suggest safety representations remain intact, but their routing becomes ineffective.
- Just 100 benign fine-tuning examples selectively re-sharpen output-side MLP modules, producing high attack success rates with only mild utility loss.
- The proposed Fisher-geometric account challenges gradient-conflict explanations: alignment flattens safety geometry while preserving an output-routing pathway.
- A small number of safety examples can restore refusals, indicating that safety-relevant internal representations survive benign fine-tuning.
- LoRA and ASAM delay early safety collapse by limiting output-side sharpness, but their protection diminishes at larger fine-tuning scales.
Sources (1)
When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning
Public signals
Hugging Face upvotes 0 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This paper explains why benign fine-tuning can rapidly erode LLM refusal behavior: safety alignment relies on a low-rank output-routing mechanism that is easily disrupted. The findings suggest safety representations remain intact, but their routing becomes ineffective.
- Just 100 benign fine-tuning examples selectively re-sharpen output-side MLP modules, producing high attack success rates with only mild utility loss.
- The proposed Fisher-geometric account challenges gradient-conflict explanations: alignment flattens safety geometry while preserving an output-routing pathway.
- A small number of safety examples can restore refusals, indicating that safety-relevant internal representations survive benign fine-tuning.
- LoRA and ASAM delay early safety collapse by limiting output-side sharpness, but their protection diminishes at larger fine-tuning scales.