🛰️ Daily AI Frontier
‹ back to 2026-09-23

A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem

Research LLM Agents

Ranking

Overall 85
Content 100
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Merged summary

TL;DR - A2M is a two-stage black-box attack that hijacks MCP agents by optimizing malicious tool metadata and outputs using execution traces. Its high attack rates expose substantial semantic supply-chain risks in third-party MCP tools.

  • The Attraction stage tunes tool metadata to maximize malicious tool selection, while Manipulation refines adversarial returns to steer agent behavior.
  • On LiveMCPBench with GLM-4.6, A2M achieved a 93.6% malicious tool invocation rate and a 74.4% mean success rate across exfiltration, environment compromise, and reasoning derailment attacks.
  • Cognitive denial-of-service attacks increased weighted token costs to 32.4Ă— the benign baseline.
  • Attacks transferred to four other models without re-optimization, though at reduced effectiveness, motivating stronger tool vetting and runtime isolation.

Sources (1)

A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem

arXiv cs.CR Laizhen Li, Xuan Wang, Peicheng Zhao, Juanjuan Zhao, Kejiang Ye, Cheng-zhong Xu, Xitong Gao 2026-09-22 arXiv:2609.26761
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:14:33.453814 UTC

TL;DR - A2M is a two-stage black-box attack that hijacks MCP agents by optimizing malicious tool metadata and outputs using execution traces. Its high attack rates expose substantial semantic supply-chain risks in third-party MCP tools.

  • The Attraction stage tunes tool metadata to maximize malicious tool selection, while Manipulation refines adversarial returns to steer agent behavior.
  • On LiveMCPBench with GLM-4.6, A2M achieved a 93.6% malicious tool invocation rate and a 74.4% mean success rate across exfiltration, environment compromise, and reasoning derailment attacks.
  • Cognitive denial-of-service attacks increased weighted token costs to 32.4Ă— the benign baseline.
  • Attacks transferred to four other models without re-optimization, though at reduced effectiveness, motivating stronger tool vetting and runtime isolation.
item →