🛰️ Daily AI Frontier
‹ back to 2026-09-02

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

arXiv cs.AI LLM Agents Haoyang Yan, Min-le Su, Hangfan Zhang, Zhanhao Li, Chen Zhang, Shao Zhang, Yang Chen, Lei Bai, Shuyue Hu 2026-09-01
Representative image for Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

TL;DR - Harness-of-Harness (HoH) structures coding-agent runs into iterative planning, coding, testing, and evaluation loops for sustained autonomous software improvement. Across three benchmarks and model-harness pairs, it delivered a 52.25% average relative gain over standalone harnesses and supported a 70-plus-iteration game-development deployment.

  • HoH balances defect repair with capability growth through small, verifiable development increments.
  • It separates implementation-time tests from independent evaluation while constraining outputs rather than prescribing agent workflows.
  • The framework progressively exposes deliverables, tools, and skills, promotes reuse, and maintains versioned project histories.
  • Benchmark gains reached 82.86% after three iterations; the multi-day deployment produced a playable first-person shooter with integrated mechanics, narrative, visuals, and audio.

view merged work →