When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models
Ranking
Overall
81
Content
100
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - Hidden states in five open-weight transformers often made logical validity highly decodable even when model answers were near chance. Weak causal effects from probe-derived interventions show that encoding validity, expressing it in outputs, and causally using it are distinct capabilities.
- Validity remained strongly decodable across held-out templates, semantic domains, and inference families.
- Validity information was also highly decodable in behaviorally incorrect examples where correctness-conditioned evaluation was applicable.
- Exhaustive leave-one-out tests identified limits to how broadly these representations generalized.
- Intervening along learned validity directions produced only weak, nonspecific effects relative to random controls.
Sources (1)
When Decodability Is Not Enough: Logical Validity Representations, Behavioral Dissociation, and Causal Tests in Language Models
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - Hidden states in five open-weight transformers often made logical validity highly decodable even when model answers were near chance. Weak causal effects from probe-derived interventions show that encoding validity, expressing it in outputs, and causally using it are distinct capabilities.
- Validity remained strongly decodable across held-out templates, semantic domains, and inference families.
- Validity information was also highly decodable in behaviorally incorrect examples where correctness-conditioned evaluation was applicable.
- Exhaustive leave-one-out tests identified limits to how broadly these representations generalized.
- Intervening along learned validity directions produced only weak, nonspecific effects relative to random controls.