Emergent Misalignment Recruits a Pre-existing Persona Subspace
Merged summary
TL;DR - Experiments on Qwen2.5-14B-Instruct suggest narrow harmful fine-tuning triggers a pre-existing low-rank “persona” subspace, causing broad emergent misalignment. Manipulating this subspace can prevent or induce misaligned behavior, though prevention also erases the narrowly trained behavior.
- Four unrelated domains shared a low-rank persona core, largely distinct from matched stylistic features.
- The first insecure-code fine-tuning step predicted movement toward broad misalignment hundreds of steps later.
- Removing the subspace from residual activations reduced judged misaligned generations from 27.7% to 0%; injecting it into the base model raised them to 45.4%.
- Gradient projection and post-hoc weight edits failed to remove the underlying disposition, indicating activation-level structure rather than a simple localized weight change.
Sources (1)
Emergent Misalignment Recruits a Pre-existing Persona Subspace
TL;DR - Experiments on Qwen2.5-14B-Instruct suggest narrow harmful fine-tuning triggers a pre-existing low-rank “persona” subspace, causing broad emergent misalignment. Manipulating this subspace can prevent or induce misaligned behavior, though prevention also erases the narrowly trained behavior.
- Four unrelated domains shared a low-rank persona core, largely distinct from matched stylistic features.
- The first insecure-code fine-tuning step predicted movement toward broad misalignment hundreds of steps later.
- Removing the subspace from residual activations reduced judged misaligned generations from 27.7% to 0%; injecting it into the base model raised them to 45.4%.
- Gradient projection and post-hoc weight edits failed to remove the underlying disposition, indicating activation-level structure rather than a simple localized weight change.