Synthetic Persona Pretraining: Alignment from Token Zero
Ranking
Overall
83
Content
95
Popularity
56
Observed public metrics from 1 member.
Merged summary
TL;DR - Synthetic Persona Pretraining embeds a constitution-aligned assistant persona from the start of language-model pretraining rather than adding alignment only afterward. Experiments suggest this early intervention improves value adherence and jailbreak robustness without sacrificing capabilities.
- Adds aligned first-person reflections, generated from a normative constitution, to standard pretraining documents.
- Uses ordinary cross-entropy pretraining, followed by “persona binding” on user-assistant dialogues.
- Models up to 3B parameters trained on 500B tokens showed fewer misaligned responses in out-of-distribution moral dilemmas.
- Introducing the method only near the end of pretraining was less effective, while its advantage increased with pretraining budget.
Sources (1)
Synthetic Persona Pretraining: Alignment from Token Zero
Public signals
Hugging Face upvotes 1
TL;DR - Synthetic Persona Pretraining embeds a constitution-aligned assistant persona from the start of language-model pretraining rather than adding alignment only afterward. Experiments suggest this early intervention improves value adherence and jailbreak robustness without sacrificing capabilities.
- Adds aligned first-person reflections, generated from a normative constitution, to standard pretraining documents.
- Uses ordinary cross-entropy pretraining, followed by “persona binding” on user-assistant dialogues.
- Models up to 3B parameters trained on 500B tokens showed fewer misaligned responses in out-of-distribution moral dilemmas.
- Introducing the method only near the end of pretraining was less effective, while its advantage increased with pretraining budget.