🛰️ Daily AI Frontier
‹ back to 2026-08-26

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

Research LLM Agents

Ranking

Overall 91
Content 100
Popularity 71

Observed public metrics from 1 member.

Representative image for StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

Merged summary

TL;DR - StarHarness evolves environment-specific agent scaffolding—such as prompts, tools, skills, subagents, and loop settings—without changing model weights. It improves enterprise benchmark performance by 20–35 percentage points and generalizes across held-out tasks and model families.

  • Uses stratified task sampling based on baseline failures, with separate search, selection, and held-out evaluation sets.
  • Achieves full-benchmark gains after only 4–12 accepted harness changes per environment.
  • Transfers without re-evolution across GPT and Qwen model families.
  • Improvements stem from repaired interfaces, encoded environment conventions, and operational knowledge that reduces false diagnoses and shortens some trajectories.

Sources (1)

StarHarness: Evolving Harnesses with Stratified Search for Enterprise Environments

arXiv cs.AI Esakkivel Esakkiraja, Denis Akhiyarov, Vikas Yadav, Sai Rajeswar, Patrice Bechard, Sridhar Nemala, Sagar Davasam 2026-08-25 arXiv:2608.24804
Public signals Hugging Face upvotes 41
Providers: Hugging Face · Upvotes 41 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-25 14:28:42.288971 UTC

TL;DR - StarHarness evolves environment-specific agent scaffolding—such as prompts, tools, skills, subagents, and loop settings—without changing model weights. It improves enterprise benchmark performance by 20–35 percentage points and generalizes across held-out tasks and model families.

  • Uses stratified task sampling based on baseline failures, with separate search, selection, and held-out evaluation sets.
  • Achieves full-benchmark gains after only 4–12 accepted harness changes per environment.
  • Transfers without re-evolution across GPT and Qwen model families.
  • Improvements stem from repaired interfaces, encoded environment conventions, and operational knowledge that reduces false diagnoses and shortens some trajectories.
item →