🛰️ Daily AI Frontier
‹ back to 2026-07-23

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

Research AI Interpretability

Ranking

Overall 87
Content 95
Popularity 69

Observed public metrics from 1 member.

Merged summary

TL;DR - Reconstruction-based activation explanations can score highly while containing false claims or private codes. RECAP instead trains internal content to remain independently probe-decodable, making explanations more verifiable and resistant to adversarial lies.

  • Only ~2% of claims from a Qwen-2.5-7B verbalizer affected reconstruction, indicating that standard scores capture gist rather than factual details.
  • Standard training produced co-adapted private codes in all five synthetic-ground-truth runs.
  • RECAP eliminated these codes in sandbox models at a cost of just 0.001 nat and improved probe-decodability in Pythia-160M.
  • RECAP probes detected both ordinary and adversarially optimized lies with AUCs of 0.96 and 0.95, respectively.

Sources (1)

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

arXiv cs.AI Hiskias Dingeto 2026-07-22 arXiv:2607.20379
Public signals Hugging Face upvotes 7
Providers: Hugging Face · Upvotes 7 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-21 14:37:13.962533 UTC

TL;DR - Reconstruction-based activation explanations can score highly while containing false claims or private codes. RECAP instead trains internal content to remain independently probe-decodable, making explanations more verifiable and resistant to adversarial lies.

  • Only ~2% of claims from a Qwen-2.5-7B verbalizer affected reconstruction, indicating that standard scores capture gist rather than factual details.
  • Standard training produced co-adapted private codes in all five synthetic-ground-truth runs.
  • RECAP eliminated these codes in sandbox models at a cost of just 0.001 nat and improved probe-decodability in Pythia-160M.
  • RECAP probes detected both ordinary and adversarially optimized lies with AUCs of 0.96 and 0.95, respectively.
item →