DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents
TL;DR - DBA-Bench evaluates database-operation agents in production-like PostgreSQL environments spanning 106 scenarios. The large safety gap between automated agents and human DBAs highlights major limitations in reliable end-to-end remediation.
- The benchmark uses live workloads, persistent state, multi-source observations, and reproducible snapshot restoration.
- Evaluation measures diagnosis, recovery outcomes, and compliance with safety constraints across seven task domains.
- Across 848 automated runs, Diagnosis, Outcome, and Safe Pass rates were 32.7%, 19.6%, and 12.4%, respectively.
- The best automated baseline achieved a 17.9% Safe Pass rate versus 93.4% for human DBAs; automated performance fell to 7.6% on hard scenarios.