🛰️ Daily AI Frontier
‹ back to 2026-08-28

Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit

Research LLM Evaluation

Ranking

Overall 81
Content 100
Popularity 37

Observed public metrics from 1 member.

Merged summary

TL;DR - A pre-registered audit shows that difference-in-differences analyses on bounded rating scales can falsely suggest bias in LLM judges because censoring at scale limits creates spurious interactions. This calls into question preference effects reported without accounting for differential attenuation near rating floors or ceilings.

  • The registered learner-profile effect on scaffolding preference was null: +0.085 points (95% BCa CI: −0.167 to +0.353; p = 0.684).
  • A nominally significant +0.378 interaction (p = 0.002) was not identifiable as a genuine preference difference.
  • A zero-differential-preference construction reproduced 79–85% of that interaction using the observed severity shift and rating-scale floor alone.
  • The paper derives the censoring mechanism in closed form and shows its contribution can be estimated from an audit’s own ratings.

Sources (1)

Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit

arXiv cs.CL Shuyi Fan, Boyuan Deng, Mengyu Xu, Xinhong Xie, Chenyang Li, Hongyang Zhang 2026-08-27 arXiv:2608.27309
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-31 14:13:41.707523 UTC

TL;DR - A pre-registered audit shows that difference-in-differences analyses on bounded rating scales can falsely suggest bias in LLM judges because censoring at scale limits creates spurious interactions. This calls into question preference effects reported without accounting for differential attenuation near rating floors or ceilings.

  • The registered learner-profile effect on scaffolding preference was null: +0.085 points (95% BCa CI: −0.167 to +0.353; p = 0.684).
  • A nominally significant +0.378 interaction (p = 0.002) was not identifiable as a genuine preference difference.
  • A zero-differential-preference construction reproduced 79–85% of that interaction using the observed severity shift and rating-scale floor alone.
  • The paper derives the censoring mechanism in closed form and shows its contribution can be estimated from an audit’s own ratings.
item →