🛰️ Daily AI Frontier
‹ back to 2026-08-30

FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets

Research LLM Agents

Ranking

Overall 76
Content 90
Popularity 42

Observed public metrics from 1 member.

Merged summary

TL;DR - FaulT-Bench evaluates network-troubleshooting LLM agents on 200 scenarios containing genuine faults and unreliable user tickets. It reveals that agents perform well when tickets are accurate but often invent root causes when a reported fault does not exist.

  • Covers eight network topologies and tests false reports, incorrect device attribution, and incorrect root-cause claims alongside genuine faults.
  • An automated Kathará/NIKA harness scores SADE, ReAct, and Claude Code on diagnosis outcome, proposed fix, and reasoning quality.
  • All three agents are near-saturated on accurate tickets but degrade sharply on healthy networks, often misclassifying benign conditions as faults.
  • Ticket style strongly affects results: vague reports cause greater degradation than confidently incorrect ones, while agents exhibit different failure patterns and costs.

Sources (1)

FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets

arXiv cs.NI Kuan-Hao Tseng, Niruth Bogahawatta, Yasod Ginige, Kunjan Patel, Kosta Dakic, Suranga Seneviratne 2026-08-27 arXiv:2608.27021
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-14 14:14:29.559208 UTC

TL;DR - FaulT-Bench evaluates network-troubleshooting LLM agents on 200 scenarios containing genuine faults and unreliable user tickets. It reveals that agents perform well when tickets are accurate but often invent root causes when a reported fault does not exist.

  • Covers eight network topologies and tests false reports, incorrect device attribution, and incorrect root-cause claims alongside genuine faults.
  • An automated Kathará/NIKA harness scores SADE, ReAct, and Claude Code on diagnosis outcome, proposed fix, and reasoning quality.
  • All three agents are near-saturated on accurate tickets but degrade sharply on healthy networks, often misclassifying benign conditions as faults.
  • Ticket style strongly affects results: vague reports cause greater degradation than confidently incorrect ones, while agents exhibit different failure patterns and costs.
item →