API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
TL;DR - An audit of ChatGPT, Claude, and Gemini finds that benchmark results obtained through APIs do not reliably predict performance in deployed chatbot interfaces. This gap could distort model comparisons, purchasing decisions, and policy assessments.
- Across seven systems and nine benchmarks, APIs averaged 3.4 percentage points higher accuracy than corresponding interfaces.
- API evaluations also showed 2.1 percentage points higher test–retest agreement, indicating more consistent behavior.
- For ChatGPT, the API-to-interface gap exceeded the API-only performance difference between GPT 5.3 and GPT 5.4.
- Adjusting system prompts, sampling parameters, and reasoning settings did not reliably reproduce interface behavior or eliminate the gap.