🛰️ Daily AI Frontier
‹ back to 2026-08-23

Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking

Research LLM Agents

Ranking

Overall 87
Content 95
Popularity 69

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper shows that poisoning only 1.2% of an agent’s persistent-memory corpus can reduce accuracy from 0.850 to 0.300. Content screening fails to detect the false assertions, while provenance-weighted retrieval faces a fundamental tradeoff between blocking poison and retaining legitimate untrusted evidence.

  • A four-stage write-time screening pipeline rejected none of 360 poisoned memories, despite achieving 0.832 recall on indirect prompt injection.
  • The shipped provenance weight was statistically indistinguishable from no defense (p=0.80).
  • Stronger provenance weighting raised mixed-corpus accuracy from 0.3167 to 0.7000, but reduced evidence recall to zero and accuracy to 0.0417 when valid evidence was untrusted.
  • The authors propose bounded occupancy constraints during retrieval instead of additive provenance penalties.

Sources (1)

Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking

arXiv cs.CR Arulnidhi Karunanidhi 2026-08-21 arXiv:2608.21230
Public signals Semantic Scholar citations 1 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 1 · Influential citations 0 X · N/A Fetched 2026-09-19 14:24:30.965025 UTC

TL;DR - This paper shows that poisoning only 1.2% of an agent’s persistent-memory corpus can reduce accuracy from 0.850 to 0.300. Content screening fails to detect the false assertions, while provenance-weighted retrieval faces a fundamental tradeoff between blocking poison and retaining legitimate untrusted evidence.

  • A four-stage write-time screening pipeline rejected none of 360 poisoned memories, despite achieving 0.832 recall on indirect prompt injection.
  • The shipped provenance weight was statistically indistinguishable from no defense (p=0.80).
  • Stronger provenance weighting raised mixed-corpus accuracy from 0.3167 to 0.7000, but reduced evidence recall to zero and accuracy to 0.0417 when valid evidence was untrusted.
  • The authors propose bounded occupancy constraints during retrieval instead of additive provenance penalties.
item →