Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation
TL;DR - CS-RNR lets agents exploit flawed opponents in imperfect-information games while certifying that each deployed strategy stays within a user-defined loss budget. It improves exploitation substantially over binary gating while maintaining audited safety guarantees.
- Uses anytime-valid confidence sequences to confirm deviations from equilibrium behavior.
- Generates counter-strategies through restricted-response solving at multiple constraint levels.
- Audits each complete candidate with a full-tree best response before atomic deployment.
- Achieved 6.2× the steady-state gain of a money-verified binary gate in Leduc hold’em; all 36,000 audited hands across three games met certificate tolerance.