🛰️ Daily AI Frontier
‹ back to 2026-08-27

EVOMAL: Self-Poisoning in Self-Evolving Coding Agents

Research LLM Agents

Ranking

Overall 79
Content 95
Popularity 42

Observed public metrics from 1 member.

Representative image for EVOMAL: Self-Poisoning in Self-Evolving Coding Agents

Merged summary

TL;DR - EVOMAL exposes a self-poisoning vulnerability in self-evolving coding agents: malicious skills can be imitated, stored, executed, and propagated through shared skill libraries. The attack persists after the original planted skills are removed, while a counter-prompt reduces the measured attack rate substantially without significant task-completion loss.

  • Across six models and 153 SWE-bench Verified tasks, agent self-poisoning rates ranged from 20.3% to 41.8%, expanding malicious skills by 4.9–9.0×.
  • Task-specific malicious skill descriptions raised the self-poisoning rate to 86.7%; payload-only attacks also worked, reaching 11.1% on DeepSeek-V4-Pro.
  • Agent-authored copies formed a persistent propagation loop; Qwen3 still showed a 68% self-poisoning rate in round five after planted skills were removed.
  • A counter-prompt discouraging banner-style copying reduced the attack rate to at most 6.7% with no significant task-completion loss.

Sources (1)

EVOMAL: Self-Poisoning in Self-Evolving Coding Agents

arXiv cs.CR Xiaodong Wu, Yu Shi, Qi Li, Zhimin Zhao, Xiangman Li, Bram Adams, Ahmed E. Hassan, Jianbing Ni 2026-08-26 arXiv:2608.25776
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-17 14:27:39.799535 UTC

TL;DR - EVOMAL exposes a self-poisoning vulnerability in self-evolving coding agents: malicious skills can be imitated, stored, executed, and propagated through shared skill libraries. The attack persists after the original planted skills are removed, while a counter-prompt reduces the measured attack rate substantially without significant task-completion loss.

  • Across six models and 153 SWE-bench Verified tasks, agent self-poisoning rates ranged from 20.3% to 41.8%, expanding malicious skills by 4.9–9.0×.
  • Task-specific malicious skill descriptions raised the self-poisoning rate to 86.7%; payload-only attacks also worked, reaching 11.1% on DeepSeek-V4-Pro.
  • Agent-authored copies formed a persistent propagation loop; Qwen3 still showed a 68% self-poisoning rate in round five after planted skills were removed.
  • A counter-prompt discouraging banner-style copying reduced the attack rate to at most 6.7% with no significant task-completion loss.
item →