🛰️ Daily AI Frontier
‹ back to 2026-08-30

TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

arXiv cs.LG LLM Agents Tommaso Bendinelli, Artur Dox, Christian Holz 2026-08-27
Representative image for TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution

TL;DR - TraceBench is a simulation-based benchmark for evaluating how well LLM agents attribute time-series anomalies to altered parameters in physical dynamical systems. It enables controlled analysis of agent behavior and highlights the importance of domain context and output format.

  • Generates interpretable root-cause attribution tasks from three simulated mechanical systems.
  • Evaluates four LLM agents across controlled experimental conditions.
  • Agents benefit substantially from domain context and favor numerical console output over visualizations when exploring data.
  • Requiring agents to produce per-sample prediction scripts generally reduces performance compared with submitting predictions directly.

view merged work →