AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games
TL;DR - An arXiv preprint combining AIVAT variance reduction with anytime-valid confidence sequences so head-to-head agent evaluations in imperfect-information games can stop as soon as evidence suffices, cutting required hands by a median 74x. It matters because benchmarking LLM agents is expensive, and naive early stopping silently breaks stated confidence levels.
- AIVAT's conditional mean-zero corrections cut variance by a median 54x across 15 LLM agent configurations over 71,439 paired Heads-Up No-Limit Hold'em hands, but provides no stopping rule on its own.
- AV-AIVAT pairs those corrections with continuously monitored confidence sequences; the online value model trains only on past games so no hand scores its own correction, preserving validity under optional stopping.
- At 95% nominal level and ±1 Big Blind precision, raw outcomes need a median 74x more hands than AIVAT-corrected ones under the Asymptotic CS (AsympCS).
- Exact finite-sample certification uses the Empirical-Bernstein CS, requiring an independently justified payoff bound (established structurally for Leduc hold'em); a width floor from the bet cap and bound limits gains, with a median 1.37x stopping-time ratio in descriptive HUNL runs.