ML Mentorship: Some Q/A about RL | Eric Jang
One of my ML research mentees is following OpenAI's Spinning up in RL tutorials (thanks to the nice folks who put that guide together!). She emailed me some good questions about the basics of Reinforcement Learning, and I wanted to share some of my replies on my blog in case it helps further other student's understanding of RL.
One of my ML research mentees is following OpenAI's Spinning up in RL tutorials (thanks to the nice folks who put that guide together!). She emailed me some good questions about the basics of Reinforcement Learning, and I wanted to share some of my replies on my blog in case it helps further other student's understanding of RL. The classic Sutton and Barto diagram of a Markov Decision Process. Your “How to Understand ML Papers Quickly” blog post recommended asking ourselves “what loss supervises the output predictions” when reading ML papers. However, in SpinningUp, it mentions that…
saved by
related reading
- [2602.19362] LLMs Can Learn to Reason Via Off-Policy RLarxiv.org
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- Part 3: Intro to Policy Optimization - Spinning Up documentationspinningup.openai.com
- Deep Reinforcement Learning: Pong from Pixelskarpathy.github.io
- Reward is not the optimization target — LessWronglesswrong.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- An Updated Introduction to Reinforcement Learning | Sri's Blogsrianumakonda.com
- Part 2: Kinds of RL Algorithms - Spinning Up documentationspinningup.openai.com
- Harsh Bhatt (@harshbhatt7585) on Xx.com
- Key Papers in Deep RL - Spinning Up documentationspinningup.openai.com