🛰️ Daily AI Frontier
‹ back to 2026-08-30

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

Research LLMs & Foundation Models

Ranking

Overall 76
Content 90
Popularity 42

Observed public metrics from 1 member.

Representative image for JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

Merged summary

TL;DR - JudgeStealer is a query-efficient model-extraction framework that replicates black-box LLM judges across pointwise, pairwise, and listwise evaluation protocols. It highlights the vulnerability of proprietary judging capabilities even under restricted query budgets and representative defenses.

  • Converts queried pointwise scores into pairwise and listwise supervision without additional victim-model queries by exploiting cross-protocol agreement.
  • Selects informative queries using semantic diversity, predictive uncertainty, and potential judge biases.
  • Uses score smoothing and multi-protocol review to preserve score ordering and reduce catastrophic forgetting during surrogate adaptation.
  • Reaches up to 73.3% pointwise, 87.0% pairwise, and 71.6% listwise accuracy, outperforming existing extraction baselines across tested model scales and settings.

Sources (1)

JudgeStealer: Extracting LLM Judging Capabilities across Evaluation Protocols

arXiv cs.CL Chen Chen, Yaolin Chen, Xuehan Sun, Juan Lin, Xueluan Gong, Yuhang Zheng, Qian Wang, Kwok-Yan Lam 2026-08-27 arXiv:2608.26982
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-12 14:18:30.285370 UTC

TL;DR - JudgeStealer is a query-efficient model-extraction framework that replicates black-box LLM judges across pointwise, pairwise, and listwise evaluation protocols. It highlights the vulnerability of proprietary judging capabilities even under restricted query budgets and representative defenses.

  • Converts queried pointwise scores into pairwise and listwise supervision without additional victim-model queries by exploiting cross-protocol agreement.
  • Selects informative queries using semantic diversity, predictive uncertainty, and potential judge biases.
  • Uses score smoothing and multi-protocol review to preserve score ordering and reduce catastrophic forgetting during surrogate adaptation.
  • Reaches up to 73.3% pointwise, 87.0% pairwise, and 71.6% listwise accuracy, outperforming existing extraction baselines across tested model scales and settings.
item →