🛰️ Daily AI Frontier
‹ back to 2026-08-10

AI倒查论文100年!99.2%的顶刊都有问题…

量子位 AI Research Integrity 听雨 2026-08-10
Representative image for AI倒查论文100年!99.2%的顶刊都有问题…

TL;DR - 量子位 reports on a wave of AI-agent-driven audits of top-tier ML papers, finding that most ICML 2026 oral papers fail automated reproduction and that nearly all sampled papers contain at least one objectively verifiable error. It matters because agentic verification is collapsing the cost of post-publication review and may reshape how peer review and scientific credibility work.

  • A US research-review firm ran an AI-agent reproduction audit on all 168 ICML 2026 oral papers (July 22); of the 92 with ≥5 checkable claims, only 34 had >40% of claims reproduced and just 8 exceeded 80%. Failures included missing code files, broken dependency versions, mismatched outputs, and 4 papers relying on models now taken offline (permanently irreproducible).
  • Concrete discrepancies cited: a paper claiming to train "only 0.77% of base-model parameters" shipped a checkpoint training 6.31% (~8x off), and another published a judge-model reliability table with no such judge model or generating script in its repo.
  • A late-2025 GPT-5-based "Paper Correctness Checker" flagged an average 4.7 objective errors per paper with 99.2% of papers having ≥1 issue; math/formula errors dominated at 54.0%, ~30.8% of NeurIPS and ~23.8% of ICLR papers had ≥1 substantive error, and NeurIPS per-paper errors rose from 3.8 (2021) to 5.9 (2025), +55.3%.
  • Caveats stressed in the piece: irreproducible ≠ fraudulent, and the checker itself has 83.2% precision while missing roughly 40% of real errors, so human review remains required; related efforts include the Hugging Face/AlphaXiv "Agent Reproduction Challenge" for ICML 2026 and a chemist finding AI-flagged errors in 75- and ~100-year-old boiling-point reference data.

view merged work →