SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - A domain-expert audit of SciCode, the standard scientific-coding benchmark, found pervasive defects that wrongly rejected correct solutions; the corrected release (SciCode-Verified) shows frontier models are far stronger at scientific coding than reported, meaning the apparent score plateau was an artifact of the evaluation instrument, not model capability.
- A per-problem audit of all 65 test problems uncovered 263 defects, 192 of which suppress scores across 91% of main problems — via non-reproducible gold answers, over-tight tolerances, and self-contradictory specifications.
- 78% of the score-suppressing defects required specialized physics or math knowledge to detect, so ordinary clerical proofreading would not have surfaced them.
- Corrections add missing specifications, repair grading, and also tighten tests that were too lenient; each change is logged with justification and independently re-verified by a second domain expert.
- Re-evaluating twelve frontier model snapshots: subproblem accuracy rises from 45–60% to 84–98%, and main-problem accuracy from 9–27% to 69–92%.
Sources (1)
SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models
TL;DR - A domain-expert audit of SciCode, the standard scientific-coding benchmark, found pervasive defects that wrongly rejected correct solutions; the corrected release (SciCode-Verified) shows frontier models are far stronger at scientific coding than reported, meaning the apparent score plateau was an artifact of the evaluation instrument, not model capability.
- A per-problem audit of all 65 test problems uncovered 263 defects, 192 of which suppress scores across 91% of main problems — via non-reproducible gold answers, over-tight tolerances, and self-contradictory specifications.
- 78% of the score-suppressing defects required specialized physics or math knowledge to detect, so ordinary clerical proofreading would not have surfaced them.
- Corrections add missing specifications, repair grading, and also tighten tests that were too lenient; each change is logged with justification and independently re-verified by a second domain expert.
- Re-evaluating twelve frontier model snapshots: subproblem accuracy rises from 45–60% to 84–98%, and main-problem accuracy from 9–27% to 69–92%.