Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Merged summary
TL;DR - Reconstruction-based activation explanations can score highly while containing false claims or private codes. RECAP instead trains internal content to remain independently probe-decodable, making explanations more verifiable and resistant to adversarial lies.
- Only ~2% of claims from a Qwen-2.5-7B verbalizer affected reconstruction, indicating that standard scores capture gist rather than factual details.
- Standard training produced co-adapted private codes in all five synthetic-ground-truth runs.
- RECAP eliminated these codes in sandbox models at a cost of just 0.001 nat and improved probe-decodability in Pythia-160M.
- RECAP probes detected both ordinary and adversarially optimized lies with AUCs of 0.96 and 0.95, respectively.
Sources (1)
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
TL;DR - Reconstruction-based activation explanations can score highly while containing false claims or private codes. RECAP instead trains internal content to remain independently probe-decodable, making explanations more verifiable and resistant to adversarial lies.
- Only ~2% of claims from a Qwen-2.5-7B verbalizer affected reconstruction, indicating that standard scores capture gist rather than factual details.
- Standard training produced co-adapted private codes in all five synthetic-ground-truth runs.
- RECAP eliminated these codes in sandbox models at a cost of just 0.001 nat and improved probe-decodability in Pythia-160M.
- RECAP probes detected both ordinary and adversarially optimized lies with AUCs of 0.96 and 0.95, respectively.