AutoResearch神话破灭:大模型离真正自主科研还有多远?
Ranking
No observed public metrics; popularity remains neutral/archived.
Merged summary
TL;DR - AutoResearchEval evaluates eight agent–model combinations on 100 real-world frontier research tasks and finds that autonomous research agents primarily fail at scientific judgment and correction, not engineering execution. The central gap is a missing “metacognitive loop” that turns recognized flaws into revised experiments and conclusions.
- Researchers analyzed 800 end-to-end trajectories across seven scientific domains and identified 45 failure modes spanning planning, retrieval, experimentation, analysis, writing, and review.
- Cognitive and scientific failures accounted for 92.1% of observed failures, while engineering robustness issues accounted for 7.9%.
- “Uncorrected self-awareness” appeared in 660 of 800 trajectories (82.5%): agents often identified serious problems but submitted results without fixing them.
- Similar failures occurred across Claude Code, Codex, and Gemini CLI configurations, suggesting a shared limitation in current agentic research capabilities rather than a framework-specific issue.
Sources (1)
AutoResearch神话破灭:大模型离真正自主科研还有多远?
TL;DR - AutoResearchEval evaluates eight agent–model combinations on 100 real-world frontier research tasks and finds that autonomous research agents primarily fail at scientific judgment and correction, not engineering execution. The central gap is a missing “metacognitive loop” that turns recognized flaws into revised experiments and conclusions.
- Researchers analyzed 800 end-to-end trajectories across seven scientific domains and identified 45 failure modes spanning planning, retrieval, experimentation, analysis, writing, and review.
- Cognitive and scientific failures accounted for 92.1% of observed failures, while engineering robustness issues accounted for 7.9%.
- “Uncorrected self-awareness” appeared in 660 of 800 trajectories (82.5%): agents often identified serious problems but submitted results without fixing them.
- Similar failures occurred across Claude Code, Codex, and Gemini CLI configurations, suggesting a shared limitation in current agentic research capabilities rather than a framework-specific issue.