🛰️ Daily AI Frontier
‹ back to 2026-09-11

Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

arXiv cs.AI LLM Agents Minghao Guo, Meng Cao, Sui Zhao, Siyu Ning, Xin Wang, Haoze Zhao, Jiaxuan Yang, Haihong Hao, Mingfei Han, Shunlin Rong, Haijun Wu, Xiaodan Liang, Xiaojun Chang 2026-09-10
Representative image for Mr.LHDR: A Benchmark for Multimodal Real-World Long-Horizon Deep Research Agents

TL;DR - Mr.LHDR is a multimodal benchmark for testing deep-research agents on long, dependency-heavy evidence chains. Its results show that current systems struggle to maintain consistent reasoning across many interdependent steps, and final-answer accuracy can overstate true research success.

  • Tasks require an average of 12.1 intermediate conclusions and a mean dependency depth of 10.4.
  • The benchmark includes consequential non-text evidence such as images, maps, PDFs, charts, tables, logos, and video frames.
  • The strongest evaluated system reached 43.1% Overall Accuracy but only 34.3% Strict Accuracy.
  • Removing images lowered the Dependency-Aware Checklist Score by 12.6 points, while strict accuracy declined as reasoning chains lengthened.

view merged work →