Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Ranking
Overall
87
Content
95
Popularity
69
Observed public metrics from 1 member.
Merged summary
TL;DR - Reconstruction-based activation explanations can score highly while containing false claims or private codes. RECAP instead trains internal content to remain independently probe-decodable, making explanations more verifiable and resistant to adversarial lies.
- Only ~2% of claims from a Qwen-2.5-7B verbalizer affected reconstruction, indicating that standard scores capture gist rather than factual details.
- Standard training produced co-adapted private codes in all five synthetic-ground-truth runs.
- RECAP eliminated these codes in sandbox models at a cost of just 0.001 nat and improved probe-decodability in Pythia-160M.
- RECAP probes detected both ordinary and adversarially optimized lies with AUCs of 0.96 and 0.95, respectively.
Sources (1)
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Public signals
Hugging Face upvotes 7
TL;DR - Reconstruction-based activation explanations can score highly while containing false claims or private codes. RECAP instead trains internal content to remain independently probe-decodable, making explanations more verifiable and resistant to adversarial lies.
- Only ~2% of claims from a Qwen-2.5-7B verbalizer affected reconstruction, indicating that standard scores capture gist rather than factual details.
- Standard training produced co-adapted private codes in all five synthetic-ground-truth runs.
- RECAP eliminated these codes in sandbox models at a cost of just 0.001 nat and improved probe-decodability in Pythia-160M.
- RECAP probes detected both ordinary and adversarially optimized lies with AUCs of 0.96 and 0.95, respectively.