🛰️ Daily AI Frontier
‹ back to 2026-08-09

AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games

arXiv cs.GT Agent Evaluation Methods Boning Li, Yu Chen, Longbo Huang 2026-08-06

TL;DR - An arXiv preprint combining AIVAT variance reduction with anytime-valid confidence sequences so head-to-head agent evaluations in imperfect-information games can stop as soon as evidence suffices, cutting required hands by a median 74x. It matters because benchmarking LLM agents is expensive, and naive early stopping silently breaks stated confidence levels.

  • AIVAT's conditional mean-zero corrections cut variance by a median 54x across 15 LLM agent configurations over 71,439 paired Heads-Up No-Limit Hold'em hands, but provides no stopping rule on its own.
  • AV-AIVAT pairs those corrections with continuously monitored confidence sequences; the online value model trains only on past games so no hand scores its own correction, preserving validity under optional stopping.
  • At 95% nominal level and ±1 Big Blind precision, raw outcomes need a median 74x more hands than AIVAT-corrected ones under the Asymptotic CS (AsympCS).
  • Exact finite-sample certification uses the Empirical-Bernstein CS, requiring an independently justified payoff bound (established structurally for Leduc hold'em); a width floor from the bet cap and bound limits gains, with a median 1.37x stopping-time ratio in descriptive HUNL runs.

view merged work →