🛰️ Daily AI Frontier
‹ back to 2026-09-01

Do VLMs Share Safety Neurons Across Modalities?

Research Multimodal & Generative

Ranking

Overall 81
Content 100
Popularity 37

Observed public metrics from 1 member.

Representative image for Do VLMs Share Safety Neurons Across Modalities?

Merged summary

TL;DR - A causal analysis across 10 vision-language models finds that text-triggered refusal relies on a small, concentrated set of neurons, while visual safety signals are distributed across a much higher-dimensional subspace. This mismatch may explain why harmful requests embedded in images can bypass text-focused safety alignment.

  • Roughly 88 neurons—less than 0.01%—were associated with text safety, and targeted ablation substantially reduced refusals.
  • Ablating text-safety neurons was the only intervention that consistently reduced refusal across all tested models.
  • Text safety concentrated in about five subspace directions, whereas visual safety required at least 50.
  • The study introduces iterative ablation to account for self-repair and two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval.

Sources (1)

Do VLMs Share Safety Neurons Across Modalities?

arXiv cs.LG Jiaxuan Li, Jiahao Zhang, Duc Minh Vo, Huy H. Nguyen, Pride Kavumba, Koki Wataoka 2026-08-31 arXiv:2608.30750
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-22 14:27:48.706537 UTC

TL;DR - A causal analysis across 10 vision-language models finds that text-triggered refusal relies on a small, concentrated set of neurons, while visual safety signals are distributed across a much higher-dimensional subspace. This mismatch may explain why harmful requests embedded in images can bypass text-focused safety alignment.

  • Roughly 88 neurons—less than 0.01%—were associated with text safety, and targeted ablation substantially reduced refusals.
  • Ablating text-safety neurons was the only intervention that consistently reduced refusal across all tested models.
  • Text safety concentrated in about five subspace directions, whereas visual safety required at least 50.
  • The study introduces iterative ablation to account for self-repair and two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval.
item →