🛰️ Daily AI Frontier
‹ back to 2026-08-27

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

Research LLM Evaluation

Ranking

Overall 79
Content 95
Popularity 42

Observed public metrics from 1 member.

Merged summary

TL;DR - Prior scores embedded as metadata can anchor LLM-as-a-Judge systems, undermining the assumed independence of successive evaluations. The effect spans numerical scoring and categorical decisions, and common prompting mitigations do not eliminate it.

  • Seven of eight evaluated models showed a statistically significant overall anchoring effect across 185,271 successful evaluations, with absolute Cohen’s (d) reaching 0.71.
  • Token-level probes suggest a threshold-like response: adding anchored metadata sharply shifts score probabilities, while varying below-threshold anchor values has less impact.
  • On human-labeled industry data, anchoring prevented 48% of error corrections and flipped 10.18% of correct judgments to an assigned wrong label.
  • Chain-of-Thought and metadata-disregard warnings did not reduce the total effect, highlighting the need for model- and task-specific mitigation validation.

Sources (1)

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

arXiv cs.CL Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic 2026-08-26 arXiv:2608.25869
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-09 08:09:01.154526 UTC

TL;DR - Prior scores embedded as metadata can anchor LLM-as-a-Judge systems, undermining the assumed independence of successive evaluations. The effect spans numerical scoring and categorical decisions, and common prompting mitigations do not eliminate it.

  • Seven of eight evaluated models showed a statistically significant overall anchoring effect across 185,271 successful evaluations, with absolute Cohen’s (d) reaching 0.71.
  • Token-level probes suggest a threshold-like response: adding anchored metadata sharply shifts score probabilities, while varying below-threshold anchor values has less impact.
  • On human-labeled industry data, anchoring prevented 48% of error corrections and flipped 10.18% of correct judgments to an assigned wrong label.
  • Chain-of-Thought and metadata-disregard warnings did not reduce the total effect, highlighting the need for model- and task-specific mitigation validation.
item →