🛰️ Daily AI Frontier
‹ back to 2026-08-24

On the Threat Model of Weird Generalization and Emergent Misalignment

arXiv cs.CL LLMs & Foundation Models Miriam Wanner, Mark Dredze, William Walden 2026-08-24

TL;DR - This paper finds that weird generalization—broad behavioral changes after narrow fine-tuning—depends on fragile properties of training and evaluation data. The results suggest it is more plausible as an adversarial, deliberately engineered threat than as an inherent risk of routine fine-tuning.

  • Across three open-weight models and four datasets, dataset composition and language affected weird generalization more than dataset size.
  • Fine-tuning on data familiar from pretraining produced stronger weird generalization than novel data.
  • Measured effects varied substantially with the evaluation questions used.
  • Assessing emergent misalignment risks therefore requires careful scrutiny of both fine-tuning data and evaluation design.

view merged work →