🛰️ Daily AI Frontier
‹ back to 2026-08-28

From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench

Research AI Code Review

Ranking

Overall 78
Content 95
Popularity 37

Observed public metrics from 1 member.

Representative image for From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench

Merged summary

TL;DR - MCR-Bench evaluates LLMs on realistic, multi-round code review with defect lifecycle tracking. Results show current models struggle increasingly as reviews lengthen, exposing weaknesses in temporal alignment and long-range memory.

  • Contains 2,269 real-world review tasks across five programming languages.
  • Annotates defect descriptions, types, severity levels, and state changes across review rounds.
  • Mainstream LLMs perform poorly at both defect detection and lifecycle tracking, with performance degrading over additional rounds.
  • Semantically complex and low-salience defects are missed more often; errors are linked to cross-round temporal misalignment and inadequate long-range memory.

Sources (1)

From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench

arXiv cs.SE Dewu Zheng, Yanlin Wang, Xiwen Wang, Kefeng Duan, Hongyu Zhang, Xilin Liu, Yuchi Ma, Zibin Zheng 2026-08-27 arXiv:2608.27442
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-25 14:26:56.850890 UTC

TL;DR - MCR-Bench evaluates LLMs on realistic, multi-round code review with defect lifecycle tracking. Results show current models struggle increasingly as reviews lengthen, exposing weaknesses in temporal alignment and long-range memory.

  • Contains 2,269 real-world review tasks across five programming languages.
  • Annotates defect descriptions, types, severity levels, and state changes across review rounds.
  • Mainstream LLMs perform poorly at both defect detection and lifecycle tracking, with performance degrading over additional rounds.
  • Semantically complex and low-salience defects are missed more often; errors are linked to cross-round temporal misalignment and inadequate long-range memory.
item →