🛰️ Daily AI Frontier
‹ back to 2026-08-13

Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection

Research LLM Agents

Ranking

Overall 88
Content 100
Popularity 58

Observed public metrics from 1 member.

Merged summary

TL;DR - BENCH2ROBUST injects controlled tool failures into agent benchmarks to train and evaluate policies that retry, switch tools, or abstain. Combining runtime Bayesian Tool Memory with reinforcement learning improves recovery while preserving failure-free performance.

  • Tool failures caused a near-universal robustness gap across seven models from four families.
  • Bayesian Tool Memory improved held-out Retail robustness by up to 16.8 percentage points without retraining.
  • Curriculum-controlled reinforcement learning learned complementary recovery behaviors that remained useful without runtime memory.
  • Combining both methods achieved 40.8–45.5% performance under failure injection.

Sources (1)

Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection

arXiv cs.AI Chaoran Chen, Vy Nguyen, Ziji Zhang, Abhinav Gullapalli, Ziyi Wang, Yuxuan Lu, Dakuo Wang, Jing Huang, Zhou Yu, Jin Lai 2026-08-12 arXiv:2608.11977
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-09-09 08:16:30.464405 UTC

TL;DR - BENCH2ROBUST injects controlled tool failures into agent benchmarks to train and evaluate policies that retry, switch tools, or abstain. Combining runtime Bayesian Tool Memory with reinforcement learning improves recovery while preserving failure-free performance.

  • Tool failures caused a near-universal robustness gap across seven models from four families.
  • Bayesian Tool Memory improved held-out Retail robustness by up to 16.8 percentage points without retraining.
  • Curriculum-controlled reinforcement learning learned complementary recovery behaviors that remained useful without runtime memory.
  • Combining both methods achieved 40.8–45.5% performance under failure injection.
item →