Geopolitical Divisions Across Languages in Large Language Models
TL;DR - A study of 67,200 responses from GPT, Claude, and Gemini finds that evaluations of the war in Ukraine vary systematically across 112 prompt languages. The differences resemble real-world geopolitical divisions, suggesting that biases in multilingual training data may propagate through widely used AI systems.
- Researchers tested 20 statements about the war across 112 languages and three major model families.
- The balance of Russia-leaning versus Ukraine-leaning responses differed by language and was consistent across all three models.
- Language-grouped responses correlated with public attitudes toward Russia, UN voting patterns, and national aid to Ukraine.
- The pattern remained after removing individual statement pairs, indicating it was not driven by a single prompt formulation.