🛰️ Daily AI Frontier
‹ back to 2026-07-23

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

arXiv cs.AI AI Interpretability Hiskias Dingeto 2026-07-22

TL;DR - Reconstruction-based activation explanations can score highly while containing false claims or private codes. RECAP instead trains internal content to remain independently probe-decodable, making explanations more verifiable and resistant to adversarial lies.

  • Only ~2% of claims from a Qwen-2.5-7B verbalizer affected reconstruction, indicating that standard scores capture gist rather than factual details.
  • Standard training produced co-adapted private codes in all five synthetic-ground-truth runs.
  • RECAP eliminated these codes in sandbox models at a cost of just 0.001 nat and improved probe-decodability in Pythia-160M.
  • RECAP probes detected both ordinary and adversarially optimized lies with AUCs of 0.96 and 0.95, respectively.

view merged work →