🛰️ Daily AI Frontier
‹ back to 2026-08-24

AutoResearch神话破灭:大模型离真正自主科研还有多远?

Research LLM Agents

Ranking

Overall 85
Content 100
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for AutoResearch神话破灭:大模型离真正自主科研还有多远?

Merged summary

TL;DR - AutoResearchEval evaluates eight agent–model combinations on 100 real-world frontier research tasks and finds that autonomous research agents primarily fail at scientific judgment and correction, not engineering execution. The central gap is a missing “metacognitive loop” that turns recognized flaws into revised experiments and conclusions.

  • Researchers analyzed 800 end-to-end trajectories across seven scientific domains and identified 45 failure modes spanning planning, retrieval, experimentation, analysis, writing, and review.
  • Cognitive and scientific failures accounted for 92.1% of observed failures, while engineering robustness issues accounted for 7.9%.
  • “Uncorrected self-awareness” appeared in 660 of 800 trajectories (82.5%): agents often identified serious problems but submitted results without fixing them.
  • Similar failures occurred across Claude Code, Codex, and Gemini CLI configurations, suggesting a shared limitation in current agentic research capabilities rather than a framework-specific issue.

Sources (1)

AutoResearch神话破灭:大模型离真正自主科研还有多远?

WeChat: PaperWeekly 2026-08-21
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-23 14:19:13.741099 UTC

TL;DR - AutoResearchEval evaluates eight agent–model combinations on 100 real-world frontier research tasks and finds that autonomous research agents primarily fail at scientific judgment and correction, not engineering execution. The central gap is a missing “metacognitive loop” that turns recognized flaws into revised experiments and conclusions.

  • Researchers analyzed 800 end-to-end trajectories across seven scientific domains and identified 45 failure modes spanning planning, retrieval, experimentation, analysis, writing, and review.
  • Cognitive and scientific failures accounted for 92.1% of observed failures, while engineering robustness issues accounted for 7.9%.
  • “Uncorrected self-awareness” appeared in 660 of 800 trajectories (82.5%): agents often identified serious problems but submitted results without fixing them.
  • Similar failures occurred across Claude Code, Codex, and Gemini CLI configurations, suggesting a shared limitation in current agentic research capabilities rather than a framework-specific issue.
item →