🛰️ Daily AI Frontier
‹ back to 2026-08-02

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

Research LLM Agents

Ranking

Overall 75
Content 90
Popularity 40

Observed public metrics from 1 member.

Merged summary

TL;DR - ParliamentBench is an open-source benchmark using a social deduction game to measure LLM agents’ reasoning, persuasion, and deception. Frontier models perform strongly, but most struggle to sustain a consistent deceptive persona.

  • Evaluates 16 LLMs across 1,600 simulated matches involving model-model, model-human, and online-game comparisons.
  • Introduces metrics for social deduction, reasoning, and deceptive consistency.
  • GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus form the strongest-performing cluster.
  • Weaker models underperform random and simple algorithmic baselines; deception retention falls below 50% for most models.

Sources (1)

Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game

arXiv cs.CL Niklas Bauer, Lars Benedikt Kaesberg, Akiko Aizawa, Jan Philip Wahle, Bela Gipp, Terry Ruas 2026-07-30 arXiv:2607.28146
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-17 09:50:04.871580 UTC

TL;DR - ParliamentBench is an open-source benchmark using a social deduction game to measure LLM agents’ reasoning, persuasion, and deception. Frontier models perform strongly, but most struggle to sustain a consistent deceptive persona.

  • Evaluates 16 LLMs across 1,600 simulated matches involving model-model, model-human, and online-game comparisons.
  • Introduces metrics for social deduction, reasoning, and deceptive consistency.
  • GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus form the strongest-performing cluster.
  • Weaker models underperform random and simple algorithmic baselines; deception retention falls below 50% for most models.
item →