ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
TL;DR - ToolHazard is a framework for automatically synthesizing adversarial, stateful tool environments to evaluate indirect prompt-injection risks in LLM agents. It enables broader security testing and generates alignment data that improves robustness without reducing benign-task utility.
- Uses environment, attacker, and user simulators to create executable environments and long-horizon tasks.
- Automatically discovers viable injection points and generates environment-specific attack payloads.
- ToolHazard-Bench reveals substantial vulnerabilities, with attack effectiveness depending on injection timing and placement.
- Generated alignment data improves security on ToolHazard-Bench and AgentDojo while preserving normal task performance.