🛰️ Daily AI Frontier
‹ back to 2026-07-27

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

Research Low-Resource Translation

Merged summary

TL;DR - An LLM pipeline converts grammar books into synthetic parallel corpora for fine-tuning machine translation models. It substantially improves translation for several endangered languages, showing how static linguistic documentation can address severe data scarcity.

  • Tested on Kalamang, Tuatschin, and Mandan, spanning three language families.
  • Synthetic-data fine-tuning beat seed-data baselines in 75% of Kalamang and 59% of Tuatschin configurations.
  • Best ChrF++ gains were +8.8, +5.3, and +3.3 across the three languages.
  • A 96-configuration factorial study varied target part of speech, retrieval granularity, and sample volume to identify effective combinations and failure points.

Sources (1)

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

arXiv cs.CL Varun Ghat Ravikumar, Sina Ahmadi, Lena Jäger, Rico Sennrich 2026-07-24 arXiv:2607.22376

TL;DR - An LLM pipeline converts grammar books into synthetic parallel corpora for fine-tuning machine translation models. It substantially improves translation for several endangered languages, showing how static linguistic documentation can address severe data scarcity.

  • Tested on Kalamang, Tuatschin, and Mandan, spanning three language families.
  • Synthetic-data fine-tuning beat seed-data baselines in 75% of Kalamang and 59% of Tuatschin configurations.
  • Best ChrF++ gains were +8.8, +5.3, and +3.3 across the three languages.
  • A 96-configuration factorial study varied target part of speech, retrieval granularity, and sample volume to identify effective combinations and failure points.
item →