Simulating low-resource experiments - Mathias Müller
Low-resource machine translation has been an active area of research for years. On a high level, what many papers on low-resource MT have in common is that theysimulate low-resource scenarios.
Low-resource machine translation has been an active area of research for years. On a high level, what many papers on low-resource MT have in common is that they simulate low-resource scenarios . What I mean is that many papers take a data set for a high-resource language pair (such as DE-EN) and randomly subsample until they arrive at the desired size. As an example, consider Gu et al. (2018) . I believe I can single them out here because the paper makes very strong claims about “very low-resource machine translation”. As they state in the abstract: Our approach is able to achieve 23 BLEU on R
related reading
- Understanding In-Context Machine Translation for Low-Resource Languages: A Case Study on Manchu - ACL Anthologyaclanthology.org
- Pretraining Strategies using Monolingual and Parallel Data for Low-Resource Machine Translation - ACL Anthologyaclanthology.org
- 2025.acl-long.429.pdfaclanthology.org
- Tulun: Transparent and Adaptable Low-resource Machine Translation - ACL Anthologyaclanthology.org
- Exploring In-context Example Generation for Machine Translationaclanthology.org
- 2026.mellm-1.2.pdfaclanthology.org
- [2605.12715] Scaling Laws for Mixture Pretraining Under Data Constraintsarxiv.org
- Sporks of AGIsergeylevine.substack.com
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- 2025.acl-srw.17.pdfaclanthology.org
- Model organisms researchers should check whether high LRs defeat their model organisms — LessWronglesswrong.com