🛰️ Daily AI Frontier
‹ back to 2026-09-20

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

arXiv cs.AI LLM Agents Tengfei Shao 2026-09-17
Representative image for Language-model groups overstate consensus when replaying human deliberation on a reasoning task

TL;DR - Replayed human reasoning discussions show that belief-anchored LLM agent groups substantially overstate consensus. Their agreement also fails to predict accuracy, with reasoning-mode agents sometimes converging almost unanimously on incorrect answers.

  • Human full-consensus estimates varied from 24% to 57% depending on participation and final-state definitions.
  • After submit-based and participation-matched adjustments, LLM groups still exceeded human consensus by roughly 34–44 percentage points.
  • The gap persisted without early stopping and after removing the potentially memorizable answer.
  • Simulated agent consensus did not match collective accuracy or reliably estimate human group outcomes.

view merged work →