🛰️ Daily AI Frontier
‹ back to 2026-08-06

SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models

Research Benchmarks & Evaluation

Ranking

Overall 76
Content 90
Popularity 43

Observed public metrics from 1 member.

Merged summary

TL;DR - A domain-expert audit of SciCode, the standard scientific-coding benchmark, found pervasive defects that wrongly rejected correct solutions; the corrected release (SciCode-Verified) shows frontier models are far stronger at scientific coding than reported, meaning the apparent score plateau was an artifact of the evaluation instrument, not model capability.

  • A per-problem audit of all 65 test problems uncovered 263 defects, 192 of which suppress scores across 91% of main problems — via non-reproducible gold answers, over-tight tolerances, and self-contradictory specifications.
  • 78% of the score-suppressing defects required specialized physics or math knowledge to detect, so ordinary clerical proofreading would not have surfaced them.
  • Corrections add missing specifications, repair grading, and also tighten tests that were too lenient; each change is logged with justification and independently re-verified by a second domain expert.
  • Re-evaluating twelve frontier model snapshots: subproblem accuracy rises from 45–60% to 84–98%, and main-problem accuracy from 9–27% to 69–92%.

Sources (1)

SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models

arXiv cs.SE Sihan Hu, Lyuhan Huang, Youjin Deng, Kun Chen 2026-08-05 arXiv:2608.04975
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-03 14:31:58.086297 UTC

TL;DR - A domain-expert audit of SciCode, the standard scientific-coding benchmark, found pervasive defects that wrongly rejected correct solutions; the corrected release (SciCode-Verified) shows frontier models are far stronger at scientific coding than reported, meaning the apparent score plateau was an artifact of the evaluation instrument, not model capability.

  • A per-problem audit of all 65 test problems uncovered 263 defects, 192 of which suppress scores across 91% of main problems — via non-reproducible gold answers, over-tight tolerances, and self-contradictory specifications.
  • 78% of the score-suppressing defects required specialized physics or math knowledge to detect, so ordinary clerical proofreading would not have surfaced them.
  • Corrections add missing specifications, repair grading, and also tighten tests that were too lenient; each change is logged with justification and independently re-verified by a second domain expert.
  • Re-evaluating twelve frontier model snapshots: subproblem accuracy rises from 45–60% to 84–98%, and main-problem accuracy from 9–27% to 69–92%.
item →