[2606.06428] Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation
Abstract:Prior work has shown that large language models (LLMs) can translate unseen or low-resource languages by undergoing continued training or even by encoding a grammar book in their context. However, both methods typically overfit specific languages, with limited zero-shot transfer at test time. To translate extremely low-resource languages at scale, we argue that LLMs must acquire the meta-skill of utilizing in-context linguistic knowledge rather than memorizing specific languages. In this paper, we propose a reinforcement learning (RL) approach to unseen language translation given rich linguistic context, using a surface-level translation metric (chrF) as the reward. Empirically, despite the lightweight reward, our RL-trained models effectively extract and apply relevant linguistic information from the provided context, leading to better translations on completely unseen languages than in-context learning or supervised fine-tuning. Our analyses suggest that outcome-based RL can extend beyond conventional reasoning tasks like math and coding to serve as a recipe for language learning from context.
View PDF HTML (experimental) Abstract:Prior work has shown that large language models (LLMs) can translate unseen or low-resource languages by undergoing continued training or even by encoding a grammar book in their context. However, both methods typically overfit specific languages, with limited zero-shot transfer at test time. To translate extremely low-resource languages at scale, we argue that LLMs must acquire the meta-skill of utilizing in-context linguistic knowledge rather than memorizing specific languages. In this paper, we propose a reinforcement learning (RL) approach to unseen…
saved by
related reading
- 2026.mellm-1.2.pdfaclanthology.org
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- Exploring In-context Example Generation for Machine Translationaclanthology.org
- LLM-based Translation Inference with Iterative Bilingual Understandingaclanthology.org
- 2025.acl-long.429.pdfaclanthology.org
- Unsupervised Elicitation of Language Modelsarxiv.org
- Learning without training: The implicit dynamics of in-context learningarxiv.org
- How does in-context learning work? A framework for understanding the differences from traditional supervised learning | SAIL Blogai.stanford.edu
- radford2018improving.pdfcs.ubc.ca
- Offline RL and Large Language Models - by Sergey Levinesergeylevine.substack.com