Phantom Gains: Auditing Self-Improvement Against a Measured Null
TL;DR - This paper shows that common transition-level evaluations can falsely report LLM self-improvement because they compare noisy measurements without a measured null. Using frozen controls, it finds no reliable gains from three Qwen3-8B self-training variants, while external distillation produces statistically supported improvements.
- Auditing three rounds of rank-32 LoRA self-training identifies seven measurement failures that can invert conclusions when frozen controls are omitted.
- Single greedy decoding can manufacture apparent capability changes through artifacts such as inference batching; simple threshold corrections still yield non-zero null effects.
- A per-problem exact test using a pooled baseline and false-discovery-rate control detects no changes on held-out frozen-control replicates.
- External distillation improves problems rarely solved by the base model, whereas self-training does not and also corrupts baseline-solved problems above the measured noise floor.