🛰️ Daily AI Frontier
‹ back to 2026-09-11

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Research LLM Agents

Ranking

Overall 83
Content 100
Popularity 42

Observed public metrics from 1 member.

Representative image for Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

Merged summary

TL;DR - Mr.LHDR is a multimodal benchmark for testing deep-research agents on long, dependency-heavy evidence chains. Its results show that current systems struggle to maintain consistent reasoning across many interdependent steps, and final-answer accuracy can overstate true research success.

  • Tasks require an average of 12.1 intermediate conclusions and a mean dependency depth of 10.4.
  • The benchmark includes consequential non-text evidence such as images, maps, PDFs, charts, tables, logos, and video frames.
  • The strongest evaluated system reached 43.1% Overall Accuracy but only 34.3% Strict Accuracy.
  • Removing images lowered the Dependency-Aware Checklist Score by 12.6 points, while strict accuracy declined as reasoning chains lengthened.

Sources (1)

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

arXiv cs.AI Minghao Guo, Meng Cao, Sui Zhao, Siyu Ning, Xin Wang, Haoze Zhao, Jiaxuan Yang, Haihong Hao, Mingfei Han, Shunlin Rong, Haijun Wu, Xiaodan Liang, Xiaojun Chang 2026-09-10 arXiv:2609.11318
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:21:01.391747 UTC

TL;DR - Mr.LHDR is a multimodal benchmark for testing deep-research agents on long, dependency-heavy evidence chains. Its results show that current systems struggle to maintain consistent reasoning across many interdependent steps, and final-answer accuracy can overstate true research success.

  • Tasks require an average of 12.1 intermediate conclusions and a mean dependency depth of 10.4.
  • The benchmark includes consequential non-text evidence such as images, maps, PDFs, charts, tables, logos, and video frames.
  • The strongest evaluated system reached 43.1% Overall Accuracy but only 34.3% Strict Accuracy.
  • Removing images lowered the Dependency-Aware Checklist Score by 12.6 points, while strict accuracy declined as reasoning chains lengthened.
item →