OpenAI研究员:我们都不读论文了
Ranking
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR — An OpenAI researcher's claim that frontier labs "don't read papers anymore" (citing exaggeration and fabrication at top conferences) is paired with SAI's large-scale reproduction audit of ICML 2026 Oral papers, which found most headline claims could not be verified. It matters because it questions the credibility of peer-reviewed ML publishing while papers remain the gatekeeping credential for entering those same labs.
- SAI attempted "execution-based review" (downloading code/models/data, running experiments, comparing to the text) on all 168 ICML 2026 Orals (~0.7% acceptance from 23,918 submissions); only 104 had open code and 105 full reproductions completed.
- Only 34 of 105 reproduced >40% of their claims and just 8 exceeded 80%; median reproduction score sat at 28–30%, rising only to 42–50% after excluding unrun/aborted/hardware-infeasible experiments.
- Failure modes: broken/missing code, incomplete instructions, broken dependencies, mismatched numbers, and 4 papers depending on models that are now offline. Cited examples include a paper claiming 0.77% trainable parameters whose released checkpoint trained 6.31% (~8×), and a reliability table whose judge model was absent from the repo.
- Cost is a structural barrier: median full re-run estimated ~$8,900 on Google Cloud on-demand pricing, 17 papers over $100k, one near $2.2M — so bad work carries high reward and near-zero risk, while industry critics are accused of "pulling up the ladder" since they still hire on publication records.
Sources (1)
OpenAI研究员:我们都不读论文了
TL;DR — An OpenAI researcher's claim that frontier labs "don't read papers anymore" (citing exaggeration and fabrication at top conferences) is paired with SAI's large-scale reproduction audit of ICML 2026 Oral papers, which found most headline claims could not be verified. It matters because it questions the credibility of peer-reviewed ML publishing while papers remain the gatekeeping credential for entering those same labs.
- SAI attempted "execution-based review" (downloading code/models/data, running experiments, comparing to the text) on all 168 ICML 2026 Orals (~0.7% acceptance from 23,918 submissions); only 104 had open code and 105 full reproductions completed.
- Only 34 of 105 reproduced >40% of their claims and just 8 exceeded 80%; median reproduction score sat at 28–30%, rising only to 42–50% after excluding unrun/aborted/hardware-infeasible experiments.
- Failure modes: broken/missing code, incomplete instructions, broken dependencies, mismatched numbers, and 4 papers depending on models that are now offline. Cited examples include a paper claiming 0.77% trainable parameters whose released checkpoint trained 6.31% (~8×), and a reliability table whose judge model was absent from the repo.
- Cost is a structural barrier: median full re-run estimated ~$8,900 on Google Cloud on-demand pricing, 17 papers over $100k, one near $2.2M — so bad work carries high reward and near-zero risk, while industry critics are accused of "pulling up the ladder" since they still hire on publication records.