🛰️ Daily AI Frontier
‹ back to 2026-09-25

SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

arXiv cs.AI LLM Agents Jiajun Wu, Leixin Sun, Zihan Tan, Yitao Liu, Shuo Li, Jiaru Qian, Shanghaoran Quan, Chuangxin Zhao, Yangxu Liao, Yang Liu, Bin Chong, Guancheng Wan 2026-09-24
Representative image for SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

TL;DR - SWE-Prometheus is a 60-repository benchmark that tests whether coding agents can independently identify and implement engineering-governance improvements while preserving behavior. It distinguishes merely adding governance artifacts from producing verified, execution-backed gains.

  • The benchmark evaluates six governance dimensions using paired evidence, clean-environment probes, behavior gates, and two independent teacher ratings.
  • Across ten models on a shared 22-repository subset, mean Normalized Governance Improvement ranged from 0.0568 to 0.5760, with behavior-breakage rates from 0% to 23%.
  • A repository-blind template scored 0.272 mean NGI on a frozen ten-repository batch, but improved only Tests & CI, Quality Gates, and Documentation—not reproducibility or dependency/security governance.
  • Results support reporting improvement, behavior preservation, evidence quality, and coverage together rather than relying on a single aggregate score.

view merged work →