🛰️ Daily AI Frontier
‹ back to 2026-08-01

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

arXiv cs.SE LLM Agents Manyi Wang, Junjielong Xu, Pinjia He 2026-07-30

TL;DR - PAIChecker is a multi-agent system for detecting mismatches between pull requests and linked issues in SWE-bench-like benchmarks. Such errors affect 13.6% of analyzed SWE-bench Verified instances and can undermine evaluations of LLM software-engineering capabilities.

  • Identifies five misalignment patterns spanning eleven fine-grained scenarios.
  • Uses pattern identification, cross-agent label synthesis, and code-level validation.
  • Achieves up to 92.12% binary accuracy on SWE-Gym and 91.67% on SWE-bench Multilingual.
  • Outperforms alternatives across all four tested LLM backbones.

view merged work →