A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books
Merged summary
TL;DR - An LLM pipeline converts grammar books into synthetic parallel corpora for fine-tuning machine translation models. It substantially improves translation for several endangered languages, showing how static linguistic documentation can address severe data scarcity.
- Tested on Kalamang, Tuatschin, and Mandan, spanning three language families.
- Synthetic-data fine-tuning beat seed-data baselines in 75% of Kalamang and 59% of Tuatschin configurations.
- Best ChrF++ gains were +8.8, +5.3, and +3.3 across the three languages.
- A 96-configuration factorial study varied target part of speech, retrieval granularity, and sample volume to identify effective combinations and failure points.
Sources (1)
A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books
TL;DR - An LLM pipeline converts grammar books into synthetic parallel corpora for fine-tuning machine translation models. It substantially improves translation for several endangered languages, showing how static linguistic documentation can address severe data scarcity.
- Tested on Kalamang, Tuatschin, and Mandan, spanning three language families.
- Synthetic-data fine-tuning beat seed-data baselines in 75% of Kalamang and 59% of Tuatschin configurations.
- Best ChrF++ gains were +8.8, +5.3, and +3.3 across the three languages.
- A 96-configuration factorial study varied target part of speech, retrieval granularity, and sample volume to identify effective combinations and failure points.