🛰️ Daily AI Frontier
‹ back to 2026-09-02

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

Research LLM Agents

Ranking

Overall 85
Content 95
Popularity 63

Observed public metrics from 1 member.

Representative image for Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

Merged summary

TL;DR - Harness-of-Harness (HoH) structures coding-agent runs into iterative planning, coding, testing, and evaluation loops for sustained autonomous software improvement. Across three benchmarks and model-harness pairs, it delivered a 52.25% average relative gain over standalone harnesses and supported a 70-plus-iteration game-development deployment.

  • HoH balances defect repair with capability growth through small, verifiable development increments.
  • It separates implementation-time tests from independent evaluation while constraining outputs rather than prescribing agent workflows.
  • The framework progressively exposes deliverables, tools, and skills, promotes reuse, and maintains versioned project histories.
  • Benchmark gains reached 82.86% after three iterations; the multi-day deployment produced a playable first-person shooter with integrated mechanics, narrative, visuals, and audio.

Sources (1)

Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement

arXiv cs.AI Haoyang Yan, Min-le Su, Hangfan Zhang, Zhanhao Li, Chen Zhang, Shao Zhang, Yang Chen, Lei Bai, Shuyue Hu 2026-09-01 arXiv:2609.01481
Public signals Hugging Face upvotes 19
Providers: Hugging Face · Upvotes 19 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:24:56.315058 UTC

TL;DR - Harness-of-Harness (HoH) structures coding-agent runs into iterative planning, coding, testing, and evaluation loops for sustained autonomous software improvement. Across three benchmarks and model-harness pairs, it delivered a 52.25% average relative gain over standalone harnesses and supported a 70-plus-iteration game-development deployment.

  • HoH balances defect repair with capability growth through small, verifiable development increments.
  • It separates implementation-time tests from independent evaluation while constraining outputs rather than prescribing agent workflows.
  • The framework progressively exposes deliverables, tools, and skills, promotes reuse, and maintains versioned project histories.
  • Benchmark gains reached 82.86% after three iterations; the multi-day deployment produced a playable first-person shooter with integrated mechanics, narrative, visuals, and audio.
item →