🛰️ Daily AI Frontier
‹ back to 2026-08-30

FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets

arXiv cs.NI LLM Agents Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kunjan Patel, Kosta Dakic, Suranga Seneviratne 2026-08-27

TL;DR - FaulT-Bench evaluates network-troubleshooting LLM agents on 200 scenarios containing genuine faults and unreliable user tickets. It reveals that agents perform well when tickets are accurate but often invent root causes when a reported fault does not exist.

  • Covers eight network topologies and tests false reports, incorrect device attribution, and incorrect root-cause claims alongside genuine faults.
  • An automated Kathará/NIKA harness scores SADE, ReAct, and Claude Code on diagnosis outcome, proposed fix, and reasoning quality.
  • All three agents are near-saturated on accurate tickets but degrade sharply on healthy networks, often misclassifying benign conditions as faults.
  • Ticket style strongly affects results: vague reports cause greater degradation than confidently incorrect ones, while agents exhibit different failure patterns and costs.

view merged work →