🛰️ Daily AI Frontier
‹ back to 2026-08-07

AI能接管实验室了?中国科大最新研究给出真实物理世界的压力测试

Research LLM Agents

Ranking

Overall 65
Content 80
Popularity 32

Observed public metrics from 1 member.

Representative image for AI能接管实验室了?中国科大最新研究给出真实物理世界的压力测试

Merged summary

TL;DR - USTC built a robotic catalysis lab (45 modular workstations for synthesis, characterization, and performance testing) exposed to LLM agents as machine-readable "skills," then benchmarked whether agent-generated plans actually execute in the physical world. Only 3.3% of 4,608 trials produced workflows runnable without human repair, showing fluent planning ≠ executable science.

  • Benchmark scope: 48 configurations (6 agent frameworks × 9 LLMs) across 32 expert-defined research tasks, scored not just on plan generation but on validation, dispatch to robots, and hands-off execution.
  • Best results were Claude Code + Claude Opus 4.7 at 28.1% executable and Codex + GPT 5.5 at 19.8%; the overall unaided execution rate was 3.3% (151/4,608).
  • In a 5-round closed loop (plan → robot execution → evidence → replan), Codex/GPT 5.5 tuned recipes and conditions but never restructured its workflow skeleton or fixed persistent omissions (missing electrode binder, analyte-specific colorimetric reagent) — parameter tuning is not strategic replanning.
  • Long-horizon planning is the bottleneck: only 3 workflows exceeded 30 operation steps (max 44), and the authors frame the robotic lab as both a test bed and a future training ground, logging successes/failures as agent-alignment data.

Sources (1)

AI能接管实验室了?中国科大最新研究给出真实物理世界的压力测试

WeChat: 新智元 2026-08-06 arXiv:2607.23045
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-23 14:18:00.242513 UTC

TL;DR - USTC built a robotic catalysis lab (45 modular workstations for synthesis, characterization, and performance testing) exposed to LLM agents as machine-readable "skills," then benchmarked whether agent-generated plans actually execute in the physical world. Only 3.3% of 4,608 trials produced workflows runnable without human repair, showing fluent planning ≠ executable science.

  • Benchmark scope: 48 configurations (6 agent frameworks × 9 LLMs) across 32 expert-defined research tasks, scored not just on plan generation but on validation, dispatch to robots, and hands-off execution.
  • Best results were Claude Code + Claude Opus 4.7 at 28.1% executable and Codex + GPT 5.5 at 19.8%; the overall unaided execution rate was 3.3% (151/4,608).
  • In a 5-round closed loop (plan → robot execution → evidence → replan), Codex/GPT 5.5 tuned recipes and conditions but never restructured its workflow skeleton or fixed persistent omissions (missing electrode binder, analyte-specific colorimetric reagent) — parameter tuning is not strategic replanning.
  • Long-horizon planning is the bottleneck: only 3 workflows exceeded 30 operation steps (max 44), and the authors frame the robotic lab as both a test bed and a future training ground, logging successes/failures as agent-alignment data.
item →