🛰️ Daily AI Frontier
‹ back to 2026-08-30

TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

Research LLM Agents

Ranking

Overall 74
Content 90
Popularity 37

Observed public metrics from 1 member.

Representative image for TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

Merged summary

TL;DR - TraceBench is a simulation-based benchmark for evaluating how well LLM agents attribute time-series anomalies to altered parameters in physical dynamical systems. It enables controlled analysis of agent behavior and highlights the importance of domain context and output format.

  • Generates interpretable root-cause attribution tasks from three simulated mechanical systems.
  • Evaluates four LLM agents across controlled experimental conditions.
  • Agents benefit substantially from domain context and favor numerical console output over visualizations when exploring data.
  • Requiring agents to produce per-sample prediction scripts generally reduces performance compared with submitting predictions directly.

Sources (1)

TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

arXiv cs.LG Tommaso Bendinelli, Artur Dox, Christian Holz 2026-08-27 arXiv:2608.27182
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-21 14:28:41.286597 UTC

TL;DR - TraceBench is a simulation-based benchmark for evaluating how well LLM agents attribute time-series anomalies to altered parameters in physical dynamical systems. It enables controlled analysis of agent behavior and highlights the importance of domain context and output format.

  • Generates interpretable root-cause attribution tasks from three simulated mechanical systems.
  • Evaluates four LLM agents across controlled experimental conditions.
  • Agents benefit substantially from domain context and favor numerical console output over visualizations when exploring data.
  • Requiring agents to produce per-sample prediction scripts generally reduces performance compared with submitting predictions directly.
item →