Phantom Gains: Auditing Self-Improvement Against a Measured Null
Ranking
Overall
90
Content
100
Popularity
68
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper shows that common transition-level evaluations can falsely report LLM self-improvement because they compare noisy measurements without a measured null. Using frozen controls, it finds no reliable gains from three Qwen3-8B self-training variants, while external distillation produces statistically supported improvements.
- Auditing three rounds of rank-32 LoRA self-training identifies seven measurement failures that can invert conclusions when frozen controls are omitted.
- Single greedy decoding can manufacture apparent capability changes through artifacts such as inference batching; simple threshold corrections still yield non-zero null effects.
- A per-problem exact test using a pooled baseline and false-discovery-rate control detects no changes on held-out frozen-control replicates.
- External distillation improves problems rarely solved by the base model, whereas self-training does not and also corrupts baseline-solved problems above the measured noise floor.
Sources (1)
Phantom Gains: Auditing Self-Improvement Against a Measured Null
Public signals
Semantic Scholar citations 2 · Semantic Scholar influential citations 0
TL;DR - This paper shows that common transition-level evaluations can falsely report LLM self-improvement because they compare noisy measurements without a measured null. Using frozen controls, it finds no reliable gains from three Qwen3-8B self-training variants, while external distillation produces statistically supported improvements.
- Auditing three rounds of rank-32 LoRA self-training identifies seven measurement failures that can invert conclusions when frozen controls are omitted.
- Single greedy decoding can manufacture apparent capability changes through artifacts such as inference batching; simple threshold corrections still yield non-zero null effects.
- A per-problem exact test using a pooled baseline and false-discovery-rate control detects no changes on held-out frozen-control replicates.
- External distillation improves problems rarely solved by the base model, whereas self-training does not and also corrupts baseline-solved problems above the measured noise floor.