🛰️ Daily AI Frontier
‹ back to 2026-09-25

SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

Research LLM Agents

Ranking

Overall 85
Content 100
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

Merged summary

TL;DR - SWE-Prometheus is a 60-repository benchmark that tests whether coding agents can independently identify and implement engineering-governance improvements while preserving behavior. It distinguishes merely adding governance artifacts from producing verified, execution-backed gains.

  • The benchmark evaluates six governance dimensions using paired evidence, clean-environment probes, behavior gates, and two independent teacher ratings.
  • Across ten models on a shared 22-repository subset, mean Normalized Governance Improvement ranged from 0.0568 to 0.5760, with behavior-breakage rates from 0% to 23%.
  • A repository-blind template scored 0.272 mean NGI on a frozen ten-repository batch, but improved only Tests & CI, Quality Gates, and Documentation—not reproducibility or dependency/security governance.
  • Results support reporting improvement, behavior preservation, evidence quality, and coverage together rather than relying on a single aggregate score.

Sources (1)

SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

arXiv cs.AI Jiajun Wu, Leixin Sun, Zihan Tan, Yitao Liu, Shuo Li, Jiaru Qian, Shanghaoran Quan, Chuangxin Zhao, Yangxu Liao, Yang Liu, Bin Chong, Guancheng Wan 2026-09-24 arXiv:2609.29465
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:14:08.489371 UTC

TL;DR - SWE-Prometheus is a 60-repository benchmark that tests whether coding agents can independently identify and implement engineering-governance improvements while preserving behavior. It distinguishes merely adding governance artifacts from producing verified, execution-backed gains.

  • The benchmark evaluates six governance dimensions using paired evidence, clean-environment probes, behavior gates, and two independent teacher ratings.
  • Across ten models on a shared 22-repository subset, mean Normalized Governance Improvement ranged from 0.0568 to 0.5760, with behavior-breakage rates from 0% to 23%.
  • A repository-blind template scored 0.272 mean NGI on a frozen ten-repository batch, but improved only Tests & CI, Quality Gates, and Documentation—not reproducibility or dependency/security governance.
  • Results support reporting improvement, behavior preservation, evidence quality, and coverage together rather than relying on a single aggregate score.
item →