🛰️ Daily AI Frontier
‹ back to 2026-09-11

Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

arXiv cs.AI LLM Agents Jiaqiang Li, Yajie Yang, Zhiheng Xi, Jiadong Chen, Enyu Zhou, Senjie Jin, Yang Nan, Jiazheng Zhang, Han Wang, Yanxin Li, Dingwei Zhu, Bicheng Deng, Yuhui Wang, Xiang Zheng, Qi Zhang, Lei Bai, Xingjun Ma, Tao Gui 2026-09-10
Representative image for Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents

TL;DR - Sci-MMR is a 235-task benchmark for evaluating whether multimodal research agents can perform multi-step scientific reasoning with complete, traceable evidence. Results show that answer accuracy overstates evidence-grounded capability by more than 20 percentage points.

  • Tasks span four scientific disciplines and use argument graphs linking claims, citations, visual evidence, and supporting figure regions.
  • Evidence acquisition caused 57.2% of failures; cropping improved performance by 4.5 points, while gold evidence raised accuracy by up to 37.0 points.
  • Evidence integration caused 31.8% of failures, showing that models also struggle to derive correct conclusions from available evidence.
  • Even with gold evidence, the strongest evaluated model reached only 69.1% accuracy on the hardest tasks.

view merged work →