Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
Ranking
Overall
81
Content
100
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper introduces an uncertainty-aware LLM judging framework that chooses between parametric evaluation, web retrieval, and abstention. Calibrated thresholds provide finite-sample guarantees on the false discovery rate of accepted verdicts while improving coverage over single-mode baselines.
- Calibrates uncertainty thresholds on held-out data using Clopper–Pearson confidence intervals.
- Routes low-confidence parametric judgments to a retrieval-augmented judge for evidence-based reevaluation.
- Extends the risk guarantee to two-threshold routing without additional assumptions.
- Maintains target error rates across open-domain QA benchmarks and multiple judge scales while achieving substantially higher coverage than single-mode methods.
Sources (1)
Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This paper introduces an uncertainty-aware LLM judging framework that chooses between parametric evaluation, web retrieval, and abstention. Calibrated thresholds provide finite-sample guarantees on the false discovery rate of accepted verdicts while improving coverage over single-mode baselines.
- Calibrates uncertainty thresholds on held-out data using Clopper–Pearson confidence intervals.
- Routes low-confidence parametric judgments to a retrieval-augmented judge for evidence-based reevaluation.
- Extends the risk guarantee to two-threshold routing without additional assumptions.
- Maintains target error rates across open-domain QA benchmarks and multiple judge scales while achieving substantially higher coverage than single-mode methods.