Language-model groups overstate consensus when replaying human deliberation on a reasoning task
Ranking
Overall
75
Content
90
Popularity
39
Observed public metrics from 1 member.
Merged summary
TL;DR - Replayed human reasoning discussions show that belief-anchored LLM agent groups substantially overstate consensus. Their agreement also fails to predict accuracy, with reasoning-mode agents sometimes converging almost unanimously on incorrect answers.
- Human full-consensus estimates varied from 24% to 57% depending on participation and final-state definitions.
- After submit-based and participation-matched adjustments, LLM groups still exceeded human consensus by roughly 34–44 percentage points.
- The gap persisted without early stopping and after removing the potentially memorizable answer.
- Simulated agent consensus did not match collective accuracy or reliably estimate human group outcomes.
Sources (1)
Language-model groups overstate consensus when replaying human deliberation on a reasoning task
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - Replayed human reasoning discussions show that belief-anchored LLM agent groups substantially overstate consensus. Their agreement also fails to predict accuracy, with reasoning-mode agents sometimes converging almost unanimously on incorrect answers.
- Human full-consensus estimates varied from 24% to 57% depending on participation and final-state definitions.
- After submit-based and participation-matched adjustments, LLM groups still exceeded human consensus by roughly 34–44 percentage points.
- The gap persisted without early stopping and after removing the potentially memorizable answer.
- Simulated agent consensus did not match collective accuracy or reliably estimate human group outcomes.