TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution
TL;DR - TraceBench is a simulation-based benchmark for evaluating how well LLM agents attribute time-series anomalies to altered parameters in physical dynamical systems. It enables controlled analysis of agent behavior and highlights the importance of domain context and output format.
- Generates interpretable root-cause attribution tasks from three simulated mechanical systems.
- Evaluates four LLM agents across controlled experimental conditions.
- Agents benefit substantially from domain context and favor numerical console output over visualizations when exploring data.
- Requiring agents to produce per-sample prediction scripts generally reduces performance compared with submitting predictions directly.