HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - HarnessSafe is a benchmark of 328 executable cases that measures how attacker-influenced content persists across agent-harness state (memory, skills, tools, shared artifacts) and later hijacks a benign request. It matters because delayed, cross-session contamination is a real attack surface that end-to-end attack-success rates fail to characterize.
- Covers seven families of "persistent carriers" and is evaluated on most mainstream agent harnesses, broadening beyond prior benchmarks that test only a few carriers or a single harness.
- Each case is framed as a Persistent-Risk Lifecycle: initial attacker entry → persistence across carriers and system boundaries → later benign trigger → observable violation.
- Introduces multi-stage, trace-based evaluation that uses execution evidence to pinpoint how far an attack chain progresses and where it is contained, rather than a single pass/fail rate.
- Findings: containment is carrier-specific and depends strongly on the harness–model configuration; both harness and model backend shape outcomes, and attack-success rates obscure distinct lifecycle progression patterns.
Sources (1)
HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses
TL;DR - HarnessSafe is a benchmark of 328 executable cases that measures how attacker-influenced content persists across agent-harness state (memory, skills, tools, shared artifacts) and later hijacks a benign request. It matters because delayed, cross-session contamination is a real attack surface that end-to-end attack-success rates fail to characterize.
- Covers seven families of "persistent carriers" and is evaluated on most mainstream agent harnesses, broadening beyond prior benchmarks that test only a few carriers or a single harness.
- Each case is framed as a Persistent-Risk Lifecycle: initial attacker entry → persistence across carriers and system boundaries → later benign trigger → observable violation.
- Introduces multi-stage, trace-based evaluation that uses execution evidence to pinpoint how far an attack chain progresses and where it is contained, rather than a single pass/fail rate.
- Findings: containment is carrier-specific and depends strongly on the harness–model configuration; both harness and model backend shape outcomes, and attack-success rates obscure distinct lifecycle progression patterns.