🛰️ Daily AI Frontier
‹ back to 2026-08-08

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

arXiv cs.CL LLMs & Foundation Models Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith 2026-08-06

TL;DR - MameLoshnLM is the first open-source 8B-parameter LLM built specifically for Yiddish, released alongside a curated pretraining corpus (Oytser) and a multi-task benchmark (Kashes). It matters as a concrete template for building capable models in digitally underrepresented but textually rich languages where web-scale multilingual data fails.

  • Model is obtained by continued pretraining of Llama 3.1 8B on Oytser, a corpus mixing contemporary web-native Yiddish with literary materials.
  • Kashes is a multi-task benchmark covering translation, linguistic analysis, information extraction, and language understanding — addressing the lack of reliable Yiddish evaluation resources.
  • MameLoshnLM outperforms open baselines of similar scale across benchmark tasks, and analyses show it better captures language-defining lexical and morphological patterns than general-purpose multilingual models.
  • The authors argue existing multilingual corpora are poor Yiddish proxies due to noisy, machine-translated, and misclassified text — a failure mode they generalize to low-resource languages broadly.

view merged work →