🛰️ Daily AI Frontier
‹ back to 2026-09-09

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

Research LLMs & Foundation Models

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Representative image for API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

Merged summary

TL;DR - An audit of ChatGPT, Claude, and Gemini finds that benchmark results obtained through APIs do not reliably predict performance in deployed chatbot interfaces. This gap could distort model comparisons, purchasing decisions, and policy assessments.

  • Across seven systems and nine benchmarks, APIs averaged 3.4 percentage points higher accuracy than corresponding interfaces.
  • API evaluations also showed 2.1 percentage points higher test–retest agreement, indicating more consistent behavior.
  • For ChatGPT, the API-to-interface gap exceeded the API-only performance difference between GPT 5.3 and GPT 5.4.
  • Adjusting system prompts, sampling parameters, and reasoning settings did not reliably reproduce interface behavior or eliminate the gap.

Sources (1)

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

arXiv cs.AI Jennifer Wang, Joachim Baumann, Daniel E. Ho, Sanmi Koyejo 2026-09-08 arXiv:2609.08861
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-17 14:21:30.524076 UTC

TL;DR - An audit of ChatGPT, Claude, and Gemini finds that benchmark results obtained through APIs do not reliably predict performance in deployed chatbot interfaces. This gap could distort model comparisons, purchasing decisions, and policy assessments.

  • Across seven systems and nine benchmarks, APIs averaged 3.4 percentage points higher accuracy than corresponding interfaces.
  • API evaluations also showed 2.1 percentage points higher test–retest agreement, indicating more consistent behavior.
  • For ChatGPT, the API-to-interface gap exceeded the API-only performance difference between GPT 5.3 and GPT 5.4.
  • Adjusting system prompts, sampling parameters, and reasoning settings did not reliably reproduce interface behavior or eliminate the gap.
item →