🛰️ Daily AI Frontier
‹ back to 2026-09-02

When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

Research LLMs & Foundation Models

Ranking

Overall 80
Content 100
Popularity 34

Observed public metrics from 1 member.

Representative image for When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

Merged summary

TL;DR - This paper explains why benign fine-tuning can rapidly erode LLM refusal behavior: safety alignment relies on a low-rank output-routing mechanism that is easily disrupted. The findings suggest safety representations remain intact, but their routing becomes ineffective.

  • Just 100 benign fine-tuning examples selectively re-sharpen output-side MLP modules, producing high attack success rates with only mild utility loss.
  • The proposed Fisher-geometric account challenges gradient-conflict explanations: alignment flattens safety geometry while preserving an output-routing pathway.
  • A small number of safety examples can restore refusals, indicating that safety-relevant internal representations survive benign fine-tuning.
  • LoRA and ASAM delay early safety collapse by limiting output-side sharpness, but their protection diminishes at larger fine-tuning scales.

Sources (1)

When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning

arXiv cs.CR Yitong Guo, Xiaoyi Chen, Siyuan Zhang, Xiaofeng Wang, Haixu Tang 2026-09-01 arXiv:2609.01455
Public signals Hugging Face upvotes 0 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-25 14:24:55.197123 UTC

TL;DR - This paper explains why benign fine-tuning can rapidly erode LLM refusal behavior: safety alignment relies on a low-rank output-routing mechanism that is easily disrupted. The findings suggest safety representations remain intact, but their routing becomes ineffective.

  • Just 100 benign fine-tuning examples selectively re-sharpen output-side MLP modules, producing high attack success rates with only mild utility loss.
  • The proposed Fisher-geometric account challenges gradient-conflict explanations: alignment flattens safety geometry while preserving an output-routing pathway.
  • A small number of safety examples can restore refusals, indicating that safety-relevant internal representations survive benign fine-tuning.
  • LoRA and ASAM delay early safety collapse by limiting output-side sharpness, but their protection diminishes at larger fine-tuning scales.
item →