Joseph Suarez (e/đĄ) on X: "An Ultra Opinionated Guide to Reinforcement Learning" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Messages Home Explore Notifications Messages Grok Bookmarks Jobs Communities Premium Verified Orgs Profile More Post monkbloc @monkbloc Article See new posts Conversation Joseph Suarez (e/) @jsuarez5341 An Ultra Opinionated Guide to Reinforcement Learning 17 176 1.5K 292K Reinforcement learning is about learning through interaction. Applications include robotics, logistics, gaming, and even control problems in science like nuclear fusion. It's an underexplored niche of AI where you can really advance the field without a ton of compute. But learning RL is hard, and most of the material out there for beginners makes it even harder. The advice here is a formalized version of how I train new PufferLib contributors. Some of these came in with zero programming knowledge and now help advance our research and tools. The key is to start doing reinforcement learning immediately while filling in knowledge gaps slowly through
Joseph Suarez đĄ @jsuarez An Ultra Opinionated Guide to Reinforcement Learning 3:24 PM ¡ Jul 11, 2025 358.4K Views 18 0 1 8 177 0 1 7 7 1.9K 0 1 . 9 K 3.8K 0 3 . 8 K
Explore this link on the map ârelated reading
- Joseph Suarez đĄ on X: "https://t.co/k2z3i9Nc5p" / Xx.com
- Debugging Reinforcement Learning Systemsandyljones.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- Reformist Reinforcement Learning - by Ben Recht - arg minargmin.net
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reinforcement learning - Wikipediaen.wikipedia.org
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Deep Reinforcement Learning Doesn't Work Yetalexirpan.com
- If it makes you feel any better, I've been doing this for a while and it took me... | Hacker Newsnews.ycombinator.com
- Reward is not the optimization target â LessWronglesswrong.com
- Deep Reinforcement Learning from Human Preferencesproceedings.neurips.cc