Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
TL;DR - A pre-registered audit shows that difference-in-differences analyses on bounded rating scales can falsely suggest bias in LLM judges because censoring at scale limits creates spurious interactions. This calls into question preference effects reported without accounting for differential attenuation near rating floors or ceilings.
- The registered learner-profile effect on scaffolding preference was null: +0.085 points (95% BCa CI: −0.167 to +0.353; p = 0.684).
- A nominally significant +0.378 interaction (p = 0.002) was not identifiable as a genuine preference difference.
- A zero-differential-preference construction reproduced 79–85% of that interaction using the observed severity shift and rating-scale floor alone.
- The paper derives the censoring mechanism in closed form and shows its contribution can be estimated from an audit’s own ratings.