🛰️ Daily AI Frontier
‹ back to 2026-08-10

HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses

arXiv cs.CR LLM Agents Xiao Zhang, Yusheng Wang, Yuhao Fei, Dongyuan Li, Zian Liang, Liuyu Xiang, Hongxun Gu, Zhaofeng He 2026-08-07

TL;DR - HarnessSafe is a benchmark of 328 executable cases that measures how attacker-influenced content persists across agent-harness state (memory, skills, tools, shared artifacts) and later hijacks a benign request. It matters because delayed, cross-session contamination is a real attack surface that end-to-end attack-success rates fail to characterize.

  • Covers seven families of "persistent carriers" and is evaluated on most mainstream agent harnesses, broadening beyond prior benchmarks that test only a few carriers or a single harness.
  • Each case is framed as a Persistent-Risk Lifecycle: initial attacker entry → persistence across carriers and system boundaries → later benign trigger → observable violation.
  • Introduces multi-stage, trace-based evaluation that uses execution evidence to pinpoint how far an attack chain progresses and where it is contained, rather than a single pass/fail rate.
  • Findings: containment is carrier-specific and depends strongly on the harness–model configuration; both harness and model backend shape outcomes, and attack-success rates obscure distinct lifecycle progression patterns.

view merged work →