🛰️ Daily AI Frontier
‹ back to 2026-08-19

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

Research LLM Agents

Ranking

Overall 85
Content 95
Popularity 61

Observed public metrics from 1 member.

Representative image for HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

Merged summary

TL;DR - HarnessRisk is a 128-case benchmark for evaluating safety failures across the full lifecycle of LLM agent harnesses. Its results show that vulnerabilities depend heavily on the deployed model–harness configuration and that recognizing an attack does not reliably prevent unsafe actions.

  • Covers six phases: configuration, capability extension, runtime operation, state persistence, action control, and incident recovery.
  • Tests benign objectives paired with adversarial instructions embedded in untrusted workflow artifacts, measuring utility, attack success, persistence, and detection.
  • Across three harnesses, six models, and 14 configurations, attack success ranged from 12.6% to 80.9%, while utility remained between 75.0% and 97.6%.
  • Harness configuration was consistently the most vulnerable phase; some configurations detected risks in over 90% of runs yet still had substantial attack success.

Sources (1)

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

arXiv cs.CR Yajing Bai, Jinhao Duan, Jie Peng, Xianfeng Wu, Sijia Liu, Song Wang, Tianlong Chen 2026-08-18 arXiv:2608.17597
Public signals Hugging Face upvotes 10
Providers: Hugging Face · Upvotes 10 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-17 14:32:45.330694 UTC

TL;DR - HarnessRisk is a 128-case benchmark for evaluating safety failures across the full lifecycle of LLM agent harnesses. Its results show that vulnerabilities depend heavily on the deployed model–harness configuration and that recognizing an attack does not reliably prevent unsafe actions.

  • Covers six phases: configuration, capability extension, runtime operation, state persistence, action control, and incident recovery.
  • Tests benign objectives paired with adversarial instructions embedded in untrusted workflow artifacts, measuring utility, attack success, persistence, and detection.
  • Across three harnesses, six models, and 14 configurations, attack success ranged from 12.6% to 80.9%, while utility remained between 75.0% and 97.6%.
  • Harness configuration was consistently the most vulnerable phase; some configurations detected risks in over 90% of runs yet still had substantial attack success.
item →