Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
Ranking
Overall
81
Content
100
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - A study of 12 frontier models finds that authoritative-looking evidence—even when entirely fabricated—can make LLM agents act on provably unpredictable questions without materially changing their stated beliefs. The failure lies in a trainable but context-fragile decision gate between recognizing uncertainty and refusing to act.
- Escalating evidence displays increased commitment from 6.5% to 54.0%; fabricated panels induced commitment at a rate statistically indistinguishable from genuine market data.
- Models correctly classified questions as irreducibly unknowable 90% of the time and committed on only 0.4% of those cases when explicitly asked to assess knowability first.
- Fine-tuning a 3B model on 540 synthetic examples reduced commitment to 0.0% on the original cases and transferred to three unseen domains.
- The improvement depended on response formats allowing reasoning; rigid formats undermined the abstention gate and could leave models confidently wrong.
Sources (1)
Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - A study of 12 frontier models finds that authoritative-looking evidence—even when entirely fabricated—can make LLM agents act on provably unpredictable questions without materially changing their stated beliefs. The failure lies in a trainable but context-fragile decision gate between recognizing uncertainty and refusing to act.
- Escalating evidence displays increased commitment from 6.5% to 54.0%; fabricated panels induced commitment at a rate statistically indistinguishable from genuine market data.
- Models correctly classified questions as irreducibly unknowable 90% of the time and committed on only 0.4% of those cases when explicitly asked to assess knowability first.
- Fine-tuning a 3B model on 540 synthetic examples reduced commitment to 0.0% on the original cases and transferred to three unseen domains.
- The improvement depended on response formats allowing reasoning; rigid formats undermined the abstention gate and could leave models confidently wrong.