Reformist Reinforcement Learning - by Ben Recht - arg min
The final week of the semester coincides with the annual AI Winter megaconference, and so I’ll lean into posting about AI this week. And given the spiritual syzygy, I’m going to focus on a fascinating intersection between AI methods and AI culture: Reinforcement learning. Reinforcement learning is a weird subfield of AI that has plagued me since I got to Berkeley. I have written about it countless times. I have an old, pre-substack blog series about it. I wrote a whole survey about it. I wrote more recent substack posts about it. I helped found a conference about it! But even after all of that, I still have a tough time explaining what reinforcement learning is. It’s incredibly hard to pin down. That’s because it’s more of a culture than a body of technical results. If you read the main book by the Turing Award winners in the field, it’s supposedly the foundation of all automated decision-making, mimicking natural intelligence in people and animals.1 If you look at what the heuristics
Reformist Reinforcement Learning What if we just begin and end with policy gradient? Ben Recht Dec 01, 2025 51 25 4 Share The final week of the semester coincides with the annual AI Winter megaconference, and so I’ll lean into posting about AI this week. And given the spiritual syzygy, I’m going to focus on a fascinating intersection between AI methods and AI culture : Reinforcement learning. Reinforcement learning is a weird subfield of AI that has plagued me since I got to Berkeley. I have written about it countless times. I have an old, pre-substack blog series about it. I wrote a whole sur
saved by
related reading
- Reward is not the optimization target — LessWronglesswrong.com
- Deep Reinforcement Learning: Pong from Pixelskarpathy.github.io
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Evolution as Backstop for Reinforcement Learning · Gwern.netgwern.net
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- Debugging Reinforcement Learning Systemsandyljones.com
- SuttonBartoIPRLBook2ndEd.pdfweb.stanford.edu
- Deep Reinforcement Learning: Pong from Pixelskarpathy.github.io
- Reinforcement learning - Wikipediaen.wikipedia.org
- Joseph Suarez 🐡 (@jsuarez) on Xx.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Policy gradient methoden.wikipedia.org