一篇论文改写AI科研评价规则!中国公司拿出实践数据,双榜第一
TL;DR - A new “discovery episode” framework evaluates AI scientists on complete research cycles rather than closed-book answers, scoring hypothesis formation, experiment execution, and result interpretation. Deep Principle’s MIRA platform illustrates the approach with an agentic system connected to computation and automated laboratories.
- The framework records and assesses each decision across the scientific workflow, including failures, rigor, reproducibility, and whether findings are genuinely novel.
- MIRA uses collaborating agents, high-performance computing and automated wet-lab execution, plus reusable research memory to support closed-loop experimentation.
- The article reports that MIRA ranked first on Research Claw Benchmark and Science Agent Arena, scoring 81.1% on the latter while also achieving the lowest average task cost.
- Deep Principle says it is applying the platform to internal R&D pipelines in batteries, industrial coolants, and new-energy materials.