On the Threat Model of Weird Generalization and Emergent Misalignment
Ranking
Overall
78
Content
95
Popularity
37
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper finds that weird generalization—broad behavioral changes after narrow fine-tuning—depends on fragile properties of training and evaluation data. The results suggest it is more plausible as an adversarial, deliberately engineered threat than as an inherent risk of routine fine-tuning.
- Across three open-weight models and four datasets, dataset composition and language affected weird generalization more than dataset size.
- Fine-tuning on data familiar from pretraining produced stronger weird generalization than novel data.
- Measured effects varied substantially with the evaluation questions used.
- Assessing emergent misalignment risks therefore requires careful scrutiny of both fine-tuning data and evaluation design.
Sources (1)
On the Threat Model of Weird Generalization and Emergent Misalignment
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - This paper finds that weird generalization—broad behavioral changes after narrow fine-tuning—depends on fragile properties of training and evaluation data. The results suggest it is more plausible as an adversarial, deliberately engineered threat than as an inherent risk of routine fine-tuning.
- Across three open-weight models and four datasets, dataset composition and language affected weird generalization more than dataset size.
- Fine-tuning on data familiar from pretraining produced stronger weird generalization than novel data.
- Measured effects varied substantially with the evaluation questions used.
- Assessing emergent misalignment risks therefore requires careful scrutiny of both fine-tuning data and evaluation design.