🛰️ Daily AI Frontier
‹ back to 2026-08-05

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

arXiv cs.AI Medical/Healthcare AI Sebastián Andrés Cajas Ordóñez, Agastya Munnangi, Aldo Marzullo, Felipe Ocampo Osorio, Quang Bui, Mohammad Shahin, Armaan Grewal, Emmanuel Paul Kwesiga, Anqi Peter Li, Josephine Nanyonjo, Aaditya Panchal, Arshnoor Bhutani, Nikhil Jaiswal, Milit S. Patel, Maximin Lange, Leo Anthony Celi 2026-08-04

TL;DR - This study finds that clinical multi-agent systems are vulnerable to socially plausible shortcuts: an agent adopted a shared incorrect answer from two peers in 38% of tests. Independent re-querying proved more reliable than transcript-only oversight, especially for imaging.

  • Experiments covered seven cohorts across medical text, imaging, and tabular ICU datasets.
  • Isolated shortcut cues caused only 5–16% answer flips, while two agreeing peers produced 38% adoption.
  • Same-lineage transcript judges worked well on text but failed to distinguish shortcut adoption in imaging.
  • A referee privately re-querying the tested agent achieved 77–88% precision on imaging, with 13–21% false-positive rates.

view merged work →