Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents
TL;DR - Mr.LHDR is a multimodal benchmark for testing deep-research agents on long, dependency-heavy evidence chains. Its results show that current systems struggle to maintain consistent reasoning across many interdependent steps, and final-answer accuracy can overstate true research success.
- Tasks require an average of 12.1 intermediate conclusions and a mean dependency depth of 10.4.
- The benchmark includes consequential non-text evidence such as images, maps, PDFs, charts, tables, logos, and video frames.
- The strongest evaluated system reached 43.1% Overall Accuracy but only 34.3% Strict Accuracy.
- Removing images lowered the Dependency-Aware Checklist Score by 12.6 points, while strict accuracy declined as reasoning chains lengthened.