🛰️ Daily AI Frontier
‹ back to 2026-09-07

Causal evidence that language models use confidence to drive behaviour

Research LLMs & Foundation Models

Ranking

Overall 84
Content 100
Popularity 47

Observed public metrics from 1 member.

Merged summary

TL;DR - Kumaran et al. provide causal evidence that confidence signals within large language models influence whether they answer a question or abstain. This helps clarify the internal mechanisms behind selective answering and uncertainty-driven behavior.

  • The study examines LLM decisions to answer versus abstain.
  • Experimentally boosting confidence signals makes models more inclined to answer.
  • Suppressing those signals shifts behavior toward abstention.
  • The intervention supports a causal role for internal confidence, rather than a purely correlational association.

Sources (1)

Causal evidence that language models use confidence to drive behaviour

Nature Machine Intelligence Dharshan Kumaran, Nathaniel Daw, Simon Osindero, Petar Veličković, Viorica Patraucean 2026-09-07 doi:10.1038/s42256-026-01293-x
Public signals OpenAlex citations 0
Providers: Hugging Face · N/A OpenAlex · Citations 0 Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:23:33.503057 UTC

TL;DR - Kumaran et al. provide causal evidence that confidence signals within large language models influence whether they answer a question or abstain. This helps clarify the internal mechanisms behind selective answering and uncertainty-driven behavior.

  • The study examines LLM decisions to answer versus abstain.
  • Experimentally boosting confidence signals makes models more inclined to answer.
  • Suppressing those signals shifts behavior toward abstention.
  • The intervention supports a causal role for internal confidence, rather than a purely correlational association.
item →