Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
TL;DR - This study finds that clinical multi-agent systems are vulnerable to socially plausible shortcuts: an agent adopted a shared incorrect answer from two peers in 38% of tests. Independent re-querying proved more reliable than transcript-only oversight, especially for imaging.
- Experiments covered seven cohorts across medical text, imaging, and tabular ICU datasets.
- Isolated shortcut cues caused only 5–16% answer flips, while two agreeing peers produced 38% adoption.
- Same-lineage transcript judges worked well on text but failed to distinguish shortcut adoption in imaging.
- A referee privately re-querying the tested agent achieved 77–88% precision on imaging, with 13–21% false-positive rates.