🛰️ Daily AI Frontier
‹ back to 2026-08-09

大厂不再迷信顶会:Auto Research时代,论文含金量正在缩水

Industry & News AI Research Reproducibility

Ranking

Overall 61
Content 65
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for 大厂不再迷信顶会:Auto Research时代,论文含金量正在缩水

Merged summary

TL;DR - SAI ran an AI-agent-based reproduction study over all 168 ICML 2026 Oral papers, fully reproducing 105, and found most accepted top-conference results do not hold up under actual re-execution. It matters because it quantifies the gap between peer-review acceptance and verifiable science in the emerging "Auto Research" era.

  • Of 92 papers with ≥5 verifiable claims, only 8 scored above 80% reproduction and 34 above 40%; the median reproduction score was 28% (rising to 42% among the 80 papers when counting only full, non-downscaled runs).
  • 101 of 105 papers hit at least one obstacle: 58 had code that wouldn't run as released, 48 had numbers inconsistent with the paper, 42 lacked data, 38 shipped no runnable code, and 4 depended on retired/unavailable models.
  • SAI Review covered 78% of issues raised by ≥2 human reviewers and surfaced 903 code/reproducibility issues humans missed (vs. only 22 the humans caught and it missed); humans still outperformed on novelty and research positioning.
  • Full reproduction is expensive: median ~$8,900 per Oral paper at Google Cloud on-demand rates, 17 papers over $100K, top near $2.2M — a conservative estimate excluding salaries and 2–3x exploratory reruns.
  • Context: the study follows public criticism from OpenAI researcher Keller Jordan that headline-grabbing ICLR/ICML/NeurIPS papers are often oversold; SAI notes Review is still Beta and its findings need further verification.

Sources (1)

大厂不再迷信顶会:Auto Research时代,论文含金量正在缩水

WeChat: PaperWeekly 2026-08-07
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-08 14:16:09.109414 UTC

TL;DR - SAI ran an AI-agent-based reproduction study over all 168 ICML 2026 Oral papers, fully reproducing 105, and found most accepted top-conference results do not hold up under actual re-execution. It matters because it quantifies the gap between peer-review acceptance and verifiable science in the emerging "Auto Research" era.

  • Of 92 papers with ≥5 verifiable claims, only 8 scored above 80% reproduction and 34 above 40%; the median reproduction score was 28% (rising to 42% among the 80 papers when counting only full, non-downscaled runs).
  • 101 of 105 papers hit at least one obstacle: 58 had code that wouldn't run as released, 48 had numbers inconsistent with the paper, 42 lacked data, 38 shipped no runnable code, and 4 depended on retired/unavailable models.
  • SAI Review covered 78% of issues raised by ≥2 human reviewers and surfaced 903 code/reproducibility issues humans missed (vs. only 22 the humans caught and it missed); humans still outperformed on novelty and research positioning.
  • Full reproduction is expensive: median ~$8,900 per Oral paper at Google Cloud on-demand rates, 17 papers over $100K, top near $2.2M — a conservative estimate excluding salaries and 2–3x exploratory reruns.
  • Context: the study follows public criticism from OpenAI researcher Keller Jordan that headline-grabbing ICLR/ICML/NeurIPS papers are often oversold; SAI notes Review is still Beta and its findings need further verification.
item →