TSRGet The Daily Report
The Singularity Report

TSR Desk · science · 16 September 2026, 01:00 UTC

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

What
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
Who
arxiv.org
When
15 September 2026, 04:00 UTC
Category
Science
Primary source
https://arxiv.org/abs/2608.05850
What is not known
This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. It comes from a paper posted to arXiv on 15 September 2026. We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrained progress in Yiddish language modeling. Existing multilingual corpora and benchmarks are often poor proxies for the language, containing substantial amounts of noisy, machine-translated, and misclassified text. We address these gaps by introducing Oytser, a high-quality Yiddish pretraining corpus that combines contemporary web-native sources with literary materials, and Kashes, a multi-task benchmark spanning translation, linguistic analysis, information extraction, and language understanding. Using these resources, we continue pretraining Llama 3.1 8B to obtain MameLoshnLM. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages. Our results provide both a foundation for Yiddish NLP and a practical template for language model development in historically rich but digitally underrepresented languages.

Why it counts

Across the tasks in the benchmark, MameLoshnLM outperforms open baselines of similar scale. We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Our analyses show that these gains are not only quantitative: relative to general-purpose multilingual models, MameLoshnLM better captures language-defining lexical and morphological patterns, pointing to a broader failure mode of noisy web-scale multilingual data for low-resource languages.

Sources

Primary source: primary source

What is not known

This brief does not claim independent replication. Claims that appear only on X and not in the primary source stay unknown.

No clip. The article still stands.