🛰️ Daily AI Frontier
‹ back to 2026-07-31

One Human, $N$ Agents: Audit-Budget Allocation for LLM Agent Fleets under Miscalibrated, Correlated Confidence

arXiv cs.AI LLM Agents Cesare Zavattari, Alessandro Tommasi, Giuseppe Prencipe 2026-07-30

TL;DR - This paper studies how one human should allocate limited audits across many LLM agents when self-reported confidence is miscalibrated and errors are correlated. It identifies when confidence-ranked auditing becomes worse than random selection, exposing conditions under which oversight is effectively vacuous.

  • Models budgeted noisy inspection using a two-level Gaussian copula and derives a miscalibration threshold (δ^*).
  • Counterintuitively, (δ^*) increases as audit budgets shrink, while shared task difficulty drives substantial cross-family error correlation.
  • Five open-weight models exhibit near-constant, operationally unhelpful confidence; a proprietary model remains informative and below the estimated threshold.
  • Policy replays on recorded traces confirm the predicted ordering of auditing strategies.

view merged work →