From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - MCR-Bench evaluates LLMs on realistic, multi-round code review with defect lifecycle tracking. Results show current models struggle increasingly as reviews lengthen, exposing weaknesses in temporal alignment and long-range memory.
- Contains 2,269 real-world review tasks across five programming languages.
- Annotates defect descriptions, types, severity levels, and state changes across review rounds.
- Mainstream LLMs perform poorly at both defect detection and lifecycle tracking, with performance degrading over additional rounds.
- Semantically complex and low-salience defects are missed more often; errors are linked to cross-round temporal misalignment and inadequate long-range memory.
Sources (1)
From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - MCR-Bench evaluates LLMs on realistic, multi-round code review with defect lifecycle tracking. Results show current models struggle increasingly as reviews lengthen, exposing weaknesses in temporal alignment and long-range memory.
- Contains 2,269 real-world review tasks across five programming languages.
- Annotates defect descriptions, types, severity levels, and state changes across review rounds.
- Mainstream LLMs perform poorly at both defect detection and lifecycle tracking, with performance degrading over additional rounds.
- Semantically complex and low-salience defects are missed more often; errors are linked to cross-round temporal misalignment and inadequate long-range memory.