FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets
Ranking
Overall
76
Content
90
Popularity
42
Observed public metrics from 1 member.
Merged summary
TL;DR - FaulT-Bench evaluates network-troubleshooting LLM agents on 200 scenarios containing genuine faults and unreliable user tickets. It reveals that agents perform well when tickets are accurate but often invent root causes when a reported fault does not exist.
- Covers eight network topologies and tests false reports, incorrect device attribution, and incorrect root-cause claims alongside genuine faults.
- An automated Kathará/NIKA harness scores SADE, ReAct, and Claude Code on diagnosis outcome, proposed fix, and reasoning quality.
- All three agents are near-saturated on accurate tickets but degrade sharply on healthy networks, often misclassifying benign conditions as faults.
- Ticket style strongly affects results: vague reports cause greater degradation than confidently incorrect ones, while agents exhibit different failure patterns and costs.
Sources (1)
FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - FaulT-Bench evaluates network-troubleshooting LLM agents on 200 scenarios containing genuine faults and unreliable user tickets. It reveals that agents perform well when tickets are accurate but often invent root causes when a reported fault does not exist.
- Covers eight network topologies and tests false reports, incorrect device attribution, and incorrect root-cause claims alongside genuine faults.
- An automated Kathará/NIKA harness scores SADE, ReAct, and Claude Code on diagnosis outcome, proposed fix, and reasoning quality.
- All three agents are near-saturated on accurate tickets but degrade sharply on healthy networks, often misclassifying benign conditions as faults.
- Ticket style strongly affects results: vague reports cause greater degradation than confidently incorrect ones, while agents exhibit different failure patterns and costs.