🛰️ Daily AI Frontier
‹ back to 2026-08-28

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

arXiv cs.LG LLMs & Foundation Models Zhichen Dong, Zhixuan Liu, Yuyu Fan, Xiangtian Li, Shuyang Zhang, Chao Yang 2026-08-27

TL;DR - Scaling model-generated distillation datasets can amplify subtle teacher-specific traits in students, even when the training examples are off-task and never explicitly mention those traits. This creates a latent behavior-transfer risk that conventional data screening may miss.

  • Larger independent off-task datasets made an induced teacher trait more detectable in students compared with matched no-trait controls.
  • Scaling either amplified an already favored target trait or shifted behavior from a related alternative toward the intended trait.
  • Learned LoRA updates showed a parallel scaling trend, with effects observed across model families, trait types, multi-trait settings, and cross-model transfer.
  • The findings motivate trait-aware curation and evaluation of synthetic distillation data, including data that appears benign.

view merged work →