🛰️ Daily AI Frontier
‹ back to 2026-08-21

One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

arXiv cs.CL LLM Agents Zhuochun Li, Youngmin Ko, Ali Keramati, Nicola Ferri, Susana Palmaz Lopez Pelaez, Liang-Chun Tsai, Calvin Wang, Mirco Milletari, Tuhin Kundu, Vadim Smolyakov, Kjartan Olafsson, Tommy Guy 2026-08-20
Representative image for One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

TL;DR - Thinkingbox is an MCP-compatible sandbox and 507-workflow benchmark for testing agents on policy-constrained, stateful business tasks. Results reveal a major reliability gap: the best model reaches 65.36% pass@1 but succeeds across 20 attempts only 25.25% of the time.

  • Covers multi-turn workflows in retail, hospitality, insurance, banking, consulting IT, and HR support.
  • Evaluates terminal backend state with task-specific executable checks that detect missing, incorrect, or unintended side effects.
  • Provides isolated tool sessions, complete execution traces, and final-response checks for designated tasks.
  • Clean termination and valid state-changing tool calls often still produce failed outcomes, making tool-level signals poor proxies for end-to-end success.

view merged work →