Accelerating Goal-Conditioned Reinforcement Learning Algorithms and Research
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Accelerating Goal-Conditioned Reinforcement Learning Algorithms and Research Michał Bortkiewicz 1 Władek Pałucki 2 Vivek Myers 3 Tadeusz Dziarmaga 4 Tomasz Arczewski 4 Łukasz Kuciński 2,5,6 Benjamin Eysenbach 7 1 Warsaw University of Technology 2 University of Warsaw 3 UC Berkeley 4 Jagiellonian University 5 Polish Academy of Sciences 6 IDEAS NCBR 7 Princeton University michalbortkiewicz8@gmail.com wladek.palucki@gmail.com Abstract Self-supervision has the potential to transform reinforcement learning (RL), paralleling the breakthroughs it has enabled in other areas of machine learning. While
related reading
- Contrastive Learning as Goal-Conditioned Reinforcement Learningarxiv.org
- [2201.08299] Goal-Conditioned Reinforcement Learning: Problems and Solutionsarxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- A General Goal-Conditioned Minecraft Model - Pantographpantograph.com
- Interactive Visualization of RL Algorithms for LLM Trainingzcy233035.github.io
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org
- A Gallery of Methods Beyond RL — Part I: Sampling Methodsshengyu-feng.github.io
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- [2602.11399] Can We Really Learn One Representation to Optimize All Rewards?arxiv.org
- pistar06.pdfpi.website