🛰️ Daily AI Frontier
‹ back to 2026-08-19

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

Research LLMs & Foundation Models

Ranking

Overall 81
Content 100
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - This paper introduces an uncertainty-aware LLM judging framework that chooses between parametric evaluation, web retrieval, and abstention. Calibrated thresholds provide finite-sample guarantees on the false discovery rate of accepted verdicts while improving coverage over single-mode baselines.

  • Calibrates uncertainty thresholds on held-out data using Clopper–Pearson confidence intervals.
  • Routes low-confidence parametric judgments to a retrieval-augmented judge for evidence-based reevaluation.
  • Extends the risk guarantee to two-threshold routing without additional assumptions.
  • Maintains target error rates across open-domain QA benchmarks and multiple judge scales while achieving substantially higher coverage than single-mode methods.

Sources (1)

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

arXiv cs.CL Sher Badshah, Ali Emami, Hassan Sajjad 2026-08-18 arXiv:2608.17994
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-10 14:26:46.931912 UTC

TL;DR - This paper introduces an uncertainty-aware LLM judging framework that chooses between parametric evaluation, web retrieval, and abstention. Calibrated thresholds provide finite-sample guarantees on the false discovery rate of accepted verdicts while improving coverage over single-mode baselines.

  • Calibrates uncertainty thresholds on held-out data using Clopper–Pearson confidence intervals.
  • Routes low-confidence parametric judgments to a retrieval-augmented judge for evidence-based reevaluation.
  • Extends the risk guarantee to two-threshold routing without additional assumptions.
  • Maintains target error rates across open-domain QA benchmarks and multiple judge scales while achieving substantially higher coverage than single-mode methods.
item →