Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
TL;DR - Sci-MMR is a 235-task benchmark for evaluating whether multimodal research agents can perform multi-step scientific reasoning with complete, traceable evidence. Results show that answer accuracy overstates evidence-grounded capability by more than 20 percentage points.
- Tasks span four scientific disciplines and use argument graphs linking claims, citations, visual evidence, and supporting figure regions.
- Evidence acquisition caused 57.2% of failures; cropping improved performance by 4.5 points, while gold evidence raised accuracy by up to 37.0 points.
- Evidence integration caused 31.8% of failures, showing that models also struggle to derive correct conclusions from available evidence.
- Even with gold evidence, the strongest evaluated model reached only 69.1% accuracy on the hardest tasks.