Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
TL;DR - BENCH2ROBUST injects controlled tool failures into agent benchmarks to train and evaluate policies that retry, switch tools, or abstain. Combining runtime Bayesian Tool Memory with reinforcement learning improves recovery while preserving failure-free performance.
- Tool failures caused a near-universal robustness gap across seven models from four families.
- Bayesian Tool Memory improved held-out Retail robustness by up to 16.8 percentage points without retraining.
- Curriculum-controlled reinforcement learning learned complementary recovery behaviors that remained useful without runtime memory.
- Combining both methods achieved 40.8–45.5% performance under failure injection.