PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play
We introduce PopuLoRA, a population-based asymmetric self-play framework for reinforcement learning with verifiable rewards (RLVR) post-training of LLMs.
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play Authors Roger Creus Castanyer, Geoffrey Bradway, Lorenz Wolf, Maxwill Lin, Augustine N. Mavor-Parker, Matthew James Sargent Description We introduce PopuLoRA, a population-based asymmetric self-play framework for reinforcement learning with verifiable rewards (RLVR) post-training of LLMs. External Link https://arxiv.org/abs/2605.16727v1 Date May 20, 2026 Affiliations Vmax Reinforcement learning with verifiable rewards (RLVR) gives large language models (LLMs; hereafter,
saved by
related reading
- State of RL for reasoning LLMs | A. Weersaweers.de
- Composer2.pdfcursor.com
- Scaling Self-Play with Self-Guidancearxiv.org
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- 2401.10020.pdfarxiv.org
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com
- Self-Adapting Language Modelsarxiv.org
- Reinforcement Learning via Self-Distillationarxiv.org
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- [2411.00062] Evolving Alignment via Asymmetric Self-Playarxiv.org
- [2504.13837] Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?arxiv.org
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com