Do VLMs Share Safety Neurons Across Modalities?
TL;DR - A causal analysis across 10 vision-language models finds that text-triggered refusal relies on a small, concentrated set of neurons, while visual safety signals are distributed across a much higher-dimensional subspace. This mismatch may explain why harmful requests embedded in images can bypass text-focused safety alignment.
- Roughly 88 neurons—less than 0.01%—were associated with text safety, and targeted ablation substantially reduced refusals.
- Ablating text-safety neurons was the only intervention that consistently reduced refusal across all tested models.
- Text safety concentrated in about five subspace directions, whereas visual safety required at least 50.
- The study introduces iterative ablation to account for self-repair and two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval.