🛰️ Daily AI Frontier
‹ back to 2026-08-16

QuoteBench: How Matched Scores Can Hide Command-Path Failures

arXiv cs.AI LLM Agents Shangao Li, Yao Zhang, Volker Tresp, Yuanyuan Yang 2026-08-13
Representative image for QuoteBench: How Matched Scores Can Hide Command-Path Failures

TL;DR - QuoteBench shows that execution interfaces can severely distort evaluations of command-generating coding agents. Similar aggregate scores may conceal major command-path failures and model compensation.

  • Tests 56 one-shot tasks across 14 incident-derived command-quoting failure families.
  • Adding one unescaped parser reduced success by 55.4–73.2 percentage points across eight configurations.
  • Disclosing the boundary recovered 30.4–60.7 points for six configurations, but provided no benefit for two.
  • The execution path can reorder model rankings, so evaluations should report generation contracts, transport details, operating points, and final-state validators.

view merged work →