🛰️ Daily AI Frontier
‹ back to 2026-08-10

综述 | Autonomous Research Agents:AI 科学家与验证缺口

WeChat: 专知 LLM Agents 2026-08-09
Representative image for 综述 | Autonomous Research Agents:AI 科学家与验证缺口

TL;DR - A survey ("Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap") audits LLM-based autonomous research systems and finds that while capability claims have grown, the evidence needed to independently reproduce and verify their scientific claims is largely missing. It matters because it reframes AI-scientist evaluation from "can it do research?" to "can anyone check what it did?"

  • Corpus and method: 144 records → 125 deduplicated → 35 works included, with 26 entries fully coded (24 runnable systems, 2 research/position works) across 7 dimensions: lifecycle stage, autonomy level, evaluation method, released artifacts, human intervention points, novelty verification, and result-selection disclosure. Authors flag lower coder agreement on autonomy/novelty/selection, so ratios are directional audit evidence, not rankings.
  • Disclosure gap: of 24 runnable systems, 83% release code, 71% release prompts, 88% disclose at least one human intervention point — but only 38% release seeds or execution traces and only 38% report any novelty-verification method.
  • Closed-loop autonomy is mostly mechanical: among 9 systems at L4, 7 are metric-triggered/mechanical re-runs (L4-m), 1 is author-claimed, and only 1 is externally verified (L4-v) — and that one predates the LLM-agent era.
  • Proposed framing: a verification-signal ladder (formal verifiers > executable tests/process rewards > physical oracles/simulators > citation grounding > proxy metrics/human judgment > model self-judgment), plus a reviewer-facing reporting checklist binding each disclosure (code, seeds/traces, attempt counts and selection policy, baseline provenance, reviewer independence, hypothesis pre-registration) to a specific failure mode.

view merged work →