QuoteBench: How Matched Scores Can Hide Command-Path Failures
Ranking
Overall
81
Content
90
Popularity
59
Observed public metrics from 1 member.
Merged summary
TL;DR - QuoteBench shows that execution interfaces can severely distort evaluations of command-generating coding agents. Similar aggregate scores may conceal major command-path failures and model compensation.
- Tests 56 one-shot tasks across 14 incident-derived command-quoting failure families.
- Adding one unescaped parser reduced success by 55.4–73.2 percentage points across eight configurations.
- Disclosing the boundary recovered 30.4–60.7 points for six configurations, but provided no benefit for two.
- The execution path can reorder model rankings, so evaluations should report generation contracts, transport details, operating points, and final-state validators.
Sources (1)
QuoteBench: How Matched Scores Can Hide Command-Path Failures
Public signals
Hugging Face upvotes 8
TL;DR - QuoteBench shows that execution interfaces can severely distort evaluations of command-generating coding agents. Similar aggregate scores may conceal major command-path failures and model compensation.
- Tests 56 one-shot tasks across 14 incident-derived command-quoting failure families.
- Adding one unescaped parser reduced success by 55.4–73.2 percentage points across eight configurations.
- Disclosing the boundary recovered 30.4–60.7 points for six configurations, but provided no benefit for two.
- The execution path can reorder model rankings, so evaluations should report generation contracts, transport details, operating points, and final-state validators.