🛰️ Daily AI Frontier
‹ back to 2026-08-21

Phantom Gains: Auditing Self-Improvement Against a Measured Null

arXiv cs.AI LLMs & Foundation Models Cheng Xu, Nan Yan, Liming Chen, M-Tahar Kechadi 2026-08-20

TL;DR - This paper shows that common transition-level evaluations can falsely report LLM self-improvement because they compare noisy measurements without a measured null. Using frozen controls, it finds no reliable gains from three Qwen3-8B self-training variants, while external distillation produces statistically supported improvements.

  • Auditing three rounds of rank-32 LoRA self-training identifies seven measurement failures that can invert conclusions when frozen controls are omitted.
  • Single greedy decoding can manufacture apparent capability changes through artifacts such as inference batching; simple threshold corrections still yield non-zero null effects.
  • A per-problem exact test using a pooled baseline and false-discovery-rate control detects no changes on held-out frozen-control replicates.
  • External distillation improves problems rarely solved by the base model, whereas self-training does not and also corrupts baseline-solved problems above the measured noise floor.

view merged work →