Understanding R1-Zero-Like Training: A Critical Perspective
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. DeepSeek-R1-Zero has shown that reinforcement learning (RL) at scale can directly enhance the reasoning capabilities of LLMs without supervised fine-tuning. In this work, we critically examine R1-Zero-like training by analyzing its two core components: base models and RL. We investigate a wide range of base models, including DeepSeek-V3-Base, to understand how pretraining cha
Understanding R1-Zero-Like Training: A Critical Perspective Zichen Liu * † \dagger 1,2 , Changyu Chen *1,3 , Wenjun Li *3 , Penghui Qi *1,2 , Tianyu Pang 1 , Chao Du 1 , Wee Sun Lee 2 , Min Lin 1 1 Sea AI Lab 2 National University of Singapore 3 Singapore Management University ∗ Core Contributors. † Project Lead. Abstract DeepSeek-R1-Zero has shown that reinforcement learning (RL) at scale can directly enhance the reasoning capabilities of LLMs without supervised fine-tuning. In this work, we critically examine R1-Zero-like training by analyzing its two core components: base models and RL . We
Explore this link on the map →related reading
- State of RL for reasoning LLMs | A. Weersaweers.de
- DeepSeek-R1arxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- From REINFORCE to Dr. GRPOlancelqf.github.io
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyondhuggingface.co
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Composer2.pdfcursor.com
- PPO for LLMs: A Guide for Normal Peoplecameronrwolfe.substack.com