Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
Ranking
Overall
81
Content
100
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - Sci-MMR is a 235-task benchmark for evaluating whether multimodal research agents can perform multi-step scientific reasoning with complete, traceable evidence. Results show that answer accuracy overstates evidence-grounded capability by more than 20 percentage points.
- Tasks span four scientific disciplines and use argument graphs linking claims, citations, visual evidence, and supporting figure regions.
- Evidence acquisition caused 57.2% of failures; cropping improved performance by 4.5 points, while gold evidence raised accuracy by up to 37.0 points.
- Evidence integration caused 31.8% of failures, showing that models also struggle to derive correct conclusions from available evidence.
- Even with gold evidence, the strongest evaluated model reached only 69.1% accuracy on the hardest tasks.
Sources (1)
Sci-MMR: Benchmarking Multi-Step Evidence-Grounded Scientific Reasoning in Multimodal Agents
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - Sci-MMR is a 235-task benchmark for evaluating whether multimodal research agents can perform multi-step scientific reasoning with complete, traceable evidence. Results show that answer accuracy overstates evidence-grounded capability by more than 20 percentage points.
- Tasks span four scientific disciplines and use argument graphs linking claims, citations, visual evidence, and supporting figure regions.
- Evidence acquisition caused 57.2% of failures; cropping improved performance by 4.5 points, while gold evidence raised accuracy by up to 37.0 points.
- Evidence integration caused 31.8% of failures, showing that models also struggle to derive correct conclusions from available evidence.
- Even with gold evidence, the strongest evaluated model reached only 69.1% accuracy on the hardest tasks.