🛰️ Daily AI Frontier
‹ back to 2026-07-16

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

Research LLM Agents

Merged summary

TL;DR - MM-IssueLoc is a controlled benchmark that isolates repository-level issue localization using visual evidence (screenshots, error dialogs, UI states, logs), showing current LLM and retrieval systems remain unreliable at multimodal localization.

  • Contains 652 issue-PR instances across 23 languages, annotated with 7 image categories and 4 relevance levels, providing file- and function-level gold labels plus paired text-only vs. with-image evaluation.
  • Introduces VCE-based diagnostics that convert images into structured textual evidence, and a controlled multimodal retriever (MM-IssueLoc-VL-Emb).
  • Best results are weak: strongest agent hits 38.96 file Acc@5 and 22.45 function Acc@10; strongest retriever reaches 33.86 function Acc@10.
  • High scores on text-dominant SWE benchmarks don't transfer to multimodal localization, making visual evidence an explicit, testable evaluation variable.

Sources (1)

MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

arXiv cs.SE Shaoxiong Zhan, Shi Hu, Boyu Feng, Hai Lin, Andrew Gong, Zhengda Zhou, Jiaying Zhou, Yunyun Hou, Hao Su, Hai-Tao Zheng 2026-07-16 arXiv:2607.15205

TL;DR - MM-IssueLoc is a controlled benchmark that isolates repository-level issue localization using visual evidence (screenshots, error dialogs, UI states, logs), showing current LLM and retrieval systems remain unreliable at multimodal localization.

  • Contains 652 issue-PR instances across 23 languages, annotated with 7 image categories and 4 relevance levels, providing file- and function-level gold labels plus paired text-only vs. with-image evaluation.
  • Introduces VCE-based diagnostics that convert images into structured textual evidence, and a controlled multimodal retriever (MM-IssueLoc-VL-Emb).
  • Best results are weak: strongest agent hits 38.96 file Acc@5 and 22.45 function Acc@10; strongest retriever reaches 33.86 function Acc@10.
  • High scores on text-dominant SWE benchmarks don't transfer to multimodal localization, making visual evidence an explicit, testable evaluation variable.
item →