🛰️ Daily AI Frontier
‹ back to 2026-09-20

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

Research LLM Agents

Ranking

Overall 75
Content 90
Popularity 39

Observed public metrics from 1 member.

Representative image for Language-model groups overstate consensus when replaying human deliberation on a reasoning task

Merged summary

TL;DR - Replayed human reasoning discussions show that belief-anchored LLM agent groups substantially overstate consensus. Their agreement also fails to predict accuracy, with reasoning-mode agents sometimes converging almost unanimously on incorrect answers.

  • Human full-consensus estimates varied from 24% to 57% depending on participation and final-state definitions.
  • After submit-based and participation-matched adjustments, LLM groups still exceeded human consensus by roughly 34–44 percentage points.
  • The gap persisted without early stopping and after removing the potentially memorizable answer.
  • Simulated agent consensus did not match collective accuracy or reliably estimate human group outcomes.

Sources (1)

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

arXiv cs.AI Tengfei Shao 2026-09-17 arXiv:2609.20543
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-24 14:17:27.273195 UTC

TL;DR - Replayed human reasoning discussions show that belief-anchored LLM agent groups substantially overstate consensus. Their agreement also fails to predict accuracy, with reasoning-mode agents sometimes converging almost unanimously on incorrect answers.

  • Human full-consensus estimates varied from 24% to 57% depending on participation and final-state definitions.
  • After submit-based and participation-matched adjustments, LLM groups still exceeded human consensus by roughly 34–44 percentage points.
  • The gap persisted without early stopping and after removing the potentially memorizable answer.
  • Simulated agent consensus did not match collective accuracy or reliably estimate human group outcomes.
item →