🛰️ Daily AI Frontier
‹ back to 2026-08-30

When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents

Research LLM Agents

Ranking

Overall 79
Content 95
Popularity 42

Observed public metrics from 1 member.

Representative image for When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents

Merged summary

TL;DR - SARA is a runtime security framework that separates action suggestions derived from untrusted tool outputs from authorization to execute those actions. It sharply reduces agent attack success rates while preserving competitive task utility.

  • An isolated Action Probe detects action-inducing content in observations and tracks its provenance across multiple steps.
  • Tool calls are authorized only when supported by the user’s objective and evidence from previously authorized, successful executions.
  • “No-History-Promotion” prevents repeated malicious instructions from gaining authority merely by recurring in an agent’s history.
  • On AgentDojo and AgentDyn, SARA held attack success rates to at most 0.63% across four primary settings and reduced them across additional agent backbones.

Sources (1)

When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents

arXiv cs.AI Xiaokun Guo, Zhen Xu, Dongdong Huo, Yanqiu Zhang, Wei Wang, Qinfu Yang, Dongjin Yu, Yu Wang 2026-08-27 arXiv:2608.27146
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-21 14:28:42.531919 UTC

TL;DR - SARA is a runtime security framework that separates action suggestions derived from untrusted tool outputs from authorization to execute those actions. It sharply reduces agent attack success rates while preserving competitive task utility.

  • An isolated Action Probe detects action-inducing content in observations and tracks its provenance across multiple steps.
  • Tool calls are authorized only when supported by the user’s objective and evidence from previously authorized, successful executions.
  • “No-History-Promotion” prevents repeated malicious instructions from gaining authority merely by recurring in an agent’s history.
  • On AgentDojo and AgentDyn, SARA held attack success rates to at most 0.63% across four primary settings and reduced them across additional agent backbones.
item →