StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments
TL;DR - StarHarness evolves environment-specific agent scaffolding—such as prompts, tools, skills, subagents, and loop settings—without changing model weights. It improves enterprise benchmark performance by 20–35 percentage points and generalizes across held-out tasks and model families.
- Uses stratified task sampling based on baseline failures, with separate search, selection, and held-out evaluation sets.
- Achieves full-benchmark gains after only 4–12 accepted harness changes per environment.
- Transfers without re-evolution across GPT and Qwen model families.
- Improvements stem from repaired interfaces, encoded environment conventions, and operational knowledge that reduces false diagnoses and shortens some trajectories.