🛰️ Daily AI Frontier
‹ back to 2026-08-28

Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit

arXiv cs.CL LLM Evaluation Shuyi Fan, Boyuan Deng, Mengyu Xu, Xinhong Xie, Chenyang Li, Hongyang Zhang 2026-08-27

TL;DR - A pre-registered audit shows that difference-in-differences analyses on bounded rating scales can falsely suggest bias in LLM judges because censoring at scale limits creates spurious interactions. This calls into question preference effects reported without accounting for differential attenuation near rating floors or ceilings.

  • The registered learner-profile effect on scaffolding preference was null: +0.085 points (95% BCa CI: −0.167 to +0.353; p = 0.684).
  • A nominally significant +0.378 interaction (p = 0.002) was not identifiable as a genuine preference difference.
  • A zero-differential-preference construction reproduced 79–85% of that interaction using the observed severity shift and rating-scale floor alone.
  • The paper derives the censoring mechanism in closed form and shows its contribution can be estimated from an audit’s own ratings.

view merged work →