🛰️ Daily AI Frontier
‹ back to 2026-09-22

DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security

Research LLM Agents

Ranking

Overall 85
Content 100
Popularity N/A

No observed public metrics; popularity remains neutral/archived.

Representative image for DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security

Merged summary

TL;DR - DUMA-Bench evaluates LLM-agent security in dual-control settings where users and agents can both alter a shared environment. This more realistic interaction model raises attack success rates from 26.9% to 41.1%, suggesting security depends on the full user-agent-environment system.

  • Extends τ²-bench with adversarial environments spanning eight vulnerability classes, including RAG poisoning, cross-agent manipulation, and unsafe output handling.
  • Evaluates 14 models from OpenAI, Anthropic, DeepSeek, Qwen, and Z.ai across eight domains and multiple user-behavior regimes.
  • Shows that passive-user, static-control evaluations may substantially underestimate vulnerabilities in deployed agents.
  • Provides a protocol for studying security as an emergent property of interactive agent systems rather than of models alone.

Sources (1)

DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security

arXiv cs.AI Ivan Aleksandrov, German Kochnev, Sabrina Sadiekh, Yaroslav Rogoza 2026-09-21 arXiv:2609.24662
Public signals N/A
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-26 14:14:44.660323 UTC

TL;DR - DUMA-Bench evaluates LLM-agent security in dual-control settings where users and agents can both alter a shared environment. This more realistic interaction model raises attack success rates from 26.9% to 41.1%, suggesting security depends on the full user-agent-environment system.

  • Extends τ²-bench with adversarial environments spanning eight vulnerability classes, including RAG poisoning, cross-agent manipulation, and unsafe output handling.
  • Evaluates 14 models from OpenAI, Anthropic, DeepSeek, Qwen, and Z.ai across eight domains and multiple user-behavior regimes.
  • Shows that passive-user, static-control evaluations may substantially underestimate vulnerabilities in deployed agents.
  • Provides a protocol for studying security as an emergent property of interactive agent systems rather than of models alone.
item →