Causal evidence that language models use confidence to drive behaviour
TL;DR - Kumaran et al. provide causal evidence that confidence signals within large language models influence whether they answer a question or abstain. This helps clarify the internal mechanisms behind selective answering and uncertainty-driven behavior.
- The study examines LLM decisions to answer versus abstain.
- Experimentally boosting confidence signals makes models more inclined to answer.
- Suppressing those signals shifts behavior toward abstention.
- The intervention supports a causal role for internal confidence, rather than a purely correlational association.