🛰️ Daily AI Frontier
‹ back to 2026-09-07

Causal evidence that language models use confidence to drive behaviour

Nature Machine Intelligence LLMs & Foundation Models Dharshan Kumaran, Nathaniel Daw, Simon Osindero, Petar Veličković, Viorica Patraucean 2026-09-07

TL;DR - Kumaran et al. provide causal evidence that confidence signals within large language models influence whether they answer a question or abstain. This helps clarify the internal mechanisms behind selective answering and uncertainty-driven behavior.

  • The study examines LLM decisions to answer versus abstain.
  • Experimentally boosting confidence signals makes models more inclined to answer.
  • Suppressing those signals shifts behavior toward abstention.
  • The intervention supports a causal role for internal confidence, rather than a purely correlational association.

view merged work →