🛰️ Daily AI Frontier
‹ back to 2026-09-01

Do VLMs Share Safety Neurons Across Modalities?

arXiv cs.LG Multimodal & Generative Jiaxuan Li, Jiahao Zhang, Duc Minh Vo, Huy H. Nguyen, Pride Kavumba, Koki Wataoka 2026-08-31
Representative image for Do VLMs Share Safety Neurons Across Modalities?

TL;DR - A causal analysis across 10 vision-language models finds that text-triggered refusal relies on a small, concentrated set of neurons, while visual safety signals are distributed across a much higher-dimensional subspace. This mismatch may explain why harmful requests embedded in images can bypass text-focused safety alignment.

  • Roughly 88 neurons—less than 0.01%—were associated with text safety, and targeted ablation substantially reduced refusals.
  • Ablating text-safety neurons was the only intervention that consistently reduced refusal across all tested models.
  • Text safety concentrated in about five subspace directions, whereas visual safety required at least 50.
  • The study introduces iterative ablation to account for self-repair and two modality-isolated benchmarks, ViSafe-Detect and ViSafe-Eval.

view merged work →