Cross-Lingual Alignment Without Joint Training: Do Monolingual Language Models Converge on Universal Representations?
Ranking
Overall
91
Content
100
Popularity
71
Observed public metrics from 1 member.
Merged summary
TL;DR - Independently trained monolingual language models develop cross-lingually alignable internal representations without shared training data or explicit alignment objectives. This suggests language structure itself may enable modular multilingual systems assembled from monolingual models.
- Alignment strengthens with greater data and model scale, as well as closer linguistic similarity.
- A single Procrustes rotation learned from parallel sentences can map hidden states between models.
- Rotated English residual states patched into a German model transferred factual content, often changing its cloze prediction to the English donor model’s answer.
- The findings point toward model stitching, merging, and modular multilingual architectures.
Sources (1)
Cross-Lingual Alignment Without Joint Training: Do Monolingual Language Models Converge on Universal Representations?
Public signals
Semantic Scholar citations 1 · Semantic Scholar influential citations 0
TL;DR - Independently trained monolingual language models develop cross-lingually alignable internal representations without shared training data or explicit alignment objectives. This suggests language structure itself may enable modular multilingual systems assembled from monolingual models.
- Alignment strengthens with greater data and model scale, as well as closer linguistic similarity.
- A single Procrustes rotation learned from parallel sentences can map hidden states between models.
- Rotated English residual states patched into a German model transferred factual content, often changing its cloze prediction to the English donor model’s answer.
- The findings point toward model stitching, merging, and modular multilingual architectures.