Do VLMs Share Safety Neurons Across Modalities?
Ranking
Overall
81
Content
100
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - A causal analysis across 10 vision-language models finds that text-triggered refusal relies on a small, concentrated set of neurons, while visual safety signals are distributed across a much higher-dimensional subspace. This mismatch may explain why harmful requests embedded in images can bypass text-focused safety alignment.
- Roughly 88 neurons—less than 0.01%—were associated with text safety, and targeted ablation substantially reduced refusals.
- Ablating text-safety neurons was the only intervention that consistently reduced refusal across all tested models.
- Text safety concentrated in about five subspace directions, whereas visual safety required at least 50.
- The study introduces iterative ablation to account for self-repair and two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval.
Sources (1)
Do VLMs Share Safety Neurons Across Modalities?
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - A causal analysis across 10 vision-language models finds that text-triggered refusal relies on a small, concentrated set of neurons, while visual safety signals are distributed across a much higher-dimensional subspace. This mismatch may explain why harmful requests embedded in images can bypass text-focused safety alignment.
- Roughly 88 neurons—less than 0.01%—were associated with text safety, and targeted ablation substantially reduced refusals.
- Ablating text-safety neurons was the only intervention that consistently reduced refusal across all tested models.
- Text safety concentrated in about five subspace directions, whereas visual safety required at least 50.
- The study introduces iterative ablation to account for self-repair and two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval.