🛰️ Daily AI Frontier
‹ back to 2026-09-17

Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

Research Multimodal & Generative

Ranking

Overall 82
Content 95
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

Merged summary

TL;DR - Sparse autoencoder analysis suggests vision-language models often encode evidence needed to detect harmful memes but fail to route it into their final predictions. Improved calibration and targeted LoRA adaptation can recover much of this readout gap.

  • Sparse probes substantially outperformed native macro-F1 across six binary benchmarks: 0.740 vs. 0.432 for Qwen and 0.714 vs. 0.532 for Gemma.
  • Causal ablation and feature-patching experiments distinguished internally represented “silent” evidence from evidence already routed toward model outputs.
  • Calibration-only routing recovered 93.3% of the mean performance gap, while probe-distilled LoRA also improved native predictions; shared multi-task adaptation caused negative transfer.
  • Tests in Spanish and Hindi-English code-mixed settings indicated that the signal extends beyond English, requires paired visual evidence, and is not solely explained by OCR.

Sources (1)

Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

arXiv cs.CV Girish A. Koushik, Diptesh Kanojia, Helen Treharne 2026-09-16 arXiv:2609.18860
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:15:38.140173 UTC

TL;DR - Sparse autoencoder analysis suggests vision-language models often encode evidence needed to detect harmful memes but fail to route it into their final predictions. Improved calibration and targeted LoRA adaptation can recover much of this readout gap.

  • Sparse probes substantially outperformed native macro-F1 across six binary benchmarks: 0.740 vs. 0.432 for Qwen and 0.714 vs. 0.532 for Gemma.
  • Causal ablation and feature-patching experiments distinguished internally represented “silent” evidence from evidence already routed toward model outputs.
  • Calibration-only routing recovered 93.3% of the mean performance gap, while probe-distilled LoRA also improved native predictions; shared multi-task adaptation caused negative transfer.
  • Tests in Spanish and Hindi-English code-mixed settings indicated that the signal extends beyond English, requires paired visual evidence, and is not solely explained by OCR.
item →