PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play
We introduce PopuLoRA, a population-based asymmetric self-play framework for reinforcement learning with verifiable rewards (RLVR) post-training of LLMs.
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play Authors Roger Creus Castanyer, Geoffrey Bradway, Lorenz Wolf, Maxwill Lin, Augustine N. Mavor-Parker, Matthew James Sargent Description We introduce PopuLoRA, a population-based asymmetric self-play framework for reinforcement learning with verifiable rewards (RLVR) post-training of LLMs. External Link https://arxiv.org/abs/2605.16727v1 Date May 20, 2026 Affiliations Vmax Reinforcement learning with verifiable rewards (RLVR) gives large language models (LLMs; hereafter,
Explore this link on the map →saved by
related reading
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectoriesarxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- Composer2.pdfcursor.com
- DeepSeek-R1arxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- [2411.00062] Evolving Alignment via Asymmetric Self-Playarxiv.org
- Self-Adapting Language Modelsarxiv.org
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- pdfopenreview.net