Synthetic Persona Pretraining: Alignment from Token Zero
TL;DR - Synthetic Persona Pretraining embeds a constitution-aligned assistant persona from the start of language-model pretraining rather than adding alignment only afterward. Experiments suggest this early intervention improves value adherence and jailbreak robustness without sacrificing capabilities.
- Adds aligned first-person reflections, generated from a normative constitution, to standard pretraining documents.
- Uses ordinary cross-entropy pretraining, followed by “persona binding” on user-assistant dialogues.
- Models up to 3B parameters trained on 500B tokens showed fewer misaligned responses in out-of-distribution moral dilemmas.
- Introducing the method only near the end of pretraining was less effective, while its advantage increased with pretraining budget.