API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - An audit of ChatGPT, Claude, and Gemini finds that benchmark results obtained through APIs do not reliably predict performance in deployed chatbot interfaces. This gap could distort model comparisons, purchasing decisions, and policy assessments.
- Across seven systems and nine benchmarks, APIs averaged 3.4 percentage points higher accuracy than corresponding interfaces.
- API evaluations also showed 2.1 percentage points higher test–retest agreement, indicating more consistent behavior.
- For ChatGPT, the API-to-interface gap exceeded the API-only performance difference between GPT 5.3 and GPT 5.4.
- Adjusting system prompts, sampling parameters, and reasoning settings did not reliably reproduce interface behavior or eliminate the gap.
Sources (1)
API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - An audit of ChatGPT, Claude, and Gemini finds that benchmark results obtained through APIs do not reliably predict performance in deployed chatbot interfaces. This gap could distort model comparisons, purchasing decisions, and policy assessments.
- Across seven systems and nine benchmarks, APIs averaged 3.4 percentage points higher accuracy than corresponding interfaces.
- API evaluations also showed 2.1 percentage points higher test–retest agreement, indicating more consistent behavior.
- For ChatGPT, the API-to-interface gap exceeded the API-only performance difference between GPT 5.3 and GPT 5.4.
- Adjusting system prompts, sampling parameters, and reasoning settings did not reliably reproduce interface behavior or eliminate the gap.