🛰️ Daily AI Frontier
‹ back to 2026-08-10

HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

Research LLM Agents

Ranking

Overall 69
Content 80
Popularity 43

Observed public metrics from 1 member.

Merged summary

TL;DR - HarnessSafe is a benchmark of 328 executable cases that measures how attacker-influenced content persists across agent-harness state (memory, skills, tools, shared artifacts) and later hijacks a benign request. It matters because delayed, cross-session contamination is a real attack surface that end-to-end attack-success rates fail to characterize.

  • Covers seven families of "persistent carriers" and is evaluated on most mainstream agent harnesses, broadening beyond prior benchmarks that test only a few carriers or a single harness.
  • Each case is framed as a Persistent-Risk Lifecycle: initial attacker entry → persistence across carriers and system boundaries → later benign trigger → observable violation.
  • Introduces multi-stage, trace-based evaluation that uses execution evidence to pinpoint how far an attack chain progresses and where it is contained, rather than a single pass/fail rate.
  • Findings: containment is carrier-specific and depends strongly on the harness–model configuration; both harness and model backend shape outcomes, and attack-success rates obscure distinct lifecycle progression patterns.

Sources (1)

HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

arXiv cs.CR Xiao Zhang, Yusheng Wang, Yuhao Fei, Dongyuan Li, Zian Liang, Liuyu Xiang, Hongxun Gu, Zhaofeng He 2026-08-07 arXiv:2608.06984
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-17 09:45:16.334453 UTC

TL;DR - HarnessSafe is a benchmark of 328 executable cases that measures how attacker-influenced content persists across agent-harness state (memory, skills, tools, shared artifacts) and later hijacks a benign request. It matters because delayed, cross-session contamination is a real attack surface that end-to-end attack-success rates fail to characterize.

  • Covers seven families of "persistent carriers" and is evaluated on most mainstream agent harnesses, broadening beyond prior benchmarks that test only a few carriers or a single harness.
  • Each case is framed as a Persistent-Risk Lifecycle: initial attacker entry → persistence across carriers and system boundaries → later benign trigger → observable violation.
  • Introduces multi-stage, trace-based evaluation that uses execution evidence to pinpoint how far an attack chain progresses and where it is contained, rather than a single pass/fail rate.
  • Findings: containment is carrier-specific and depends strongly on the harness–model configuration; both harness and model backend shape outcomes, and attack-success rates obscure distinct lifecycle progression patterns.
item →