TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution
Ranking
Overall
74
Content
90
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - TraceBench is a simulation-based benchmark for evaluating how well LLM agents attribute time-series anomalies to altered parameters in physical dynamical systems. It enables controlled analysis of agent behavior and highlights the importance of domain context and output format.
- Generates interpretable root-cause attribution tasks from three simulated mechanical systems.
- Evaluates four LLM agents across controlled experimental conditions.
- Agents benefit substantially from domain context and favor numerical console output over visualizations when exploring data.
- Requiring agents to produce per-sample prediction scripts generally reduces performance compared with submitting predictions directly.
Sources (1)
TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - TraceBench is a simulation-based benchmark for evaluating how well LLM agents attribute time-series anomalies to altered parameters in physical dynamical systems. It enables controlled analysis of agent behavior and highlights the importance of domain context and output format.
- Generates interpretable root-cause attribution tasks from three simulated mechanical systems.
- Evaluates four LLM agents across controlled experimental conditions.
- Agents benefit substantially from domain context and favor numerical console output over visualizations when exploring data.
- Requiring agents to produce per-sample prediction scripts generally reduces performance compared with submitting predictions directly.