DUMA-Bench: A Dual-Control Multi-Agent Benchmark for Evaluating LLM Agent Security
TL;DR - DUMA-Bench evaluates LLM-agent security in dual-control settings where users and agents can both alter a shared environment. This more realistic interaction model raises attack success rates from 26.9% to 41.1%, suggesting security depends on the full user-agent-environment system.
- Extends τ²-bench with adversarial environments spanning eight vulnerability classes, including RAG poisoning, cross-agent manipulation, and unsafe output handling.
- Evaluates 14 models from OpenAI, Anthropic, DeepSeek, Qwen, and Z.ai across eight domains and multiple user-behavior regimes.
- Shows that passive-user, static-control evaluations may substantially underestimate vulnerabilities in deployed agents.
- Provides a protocol for studying security as an emergent property of interactive agent systems rather than of models alone.