Reformist Reinforcement Learning - by Ben Recht - arg min
The final week of the semester coincides with the annual AI Winter megaconference, and so I’ll lean into posting about AI this week. And given the spiritual syzygy, I’m going to focus on a fascinating intersection between AI methods and AI culture: Reinforcement learning. Reinforcement learning is a weird subfield of AI that has plagued me since I got to Berkeley. I have written about it countless times. I have an old, pre-substack blog series about it. I wrote a whole survey about it. I wrote more recent substack posts about it. I helped found a conference about it! But even after all of that, I still have a tough time explaining what reinforcement learning is. It’s incredibly hard to pin down. That’s because it’s more of a culture than a body of technical results. If you read the main book by the Turing Award winners in the field, it’s supposedly the foundation of all automated decision-making, mimicking natural intelligence in people and animals.1 If you look at what the heuristics
Reformist Reinforcement Learning What if we just begin and end with policy gradient? Ben Recht Dec 01, 2025 51 25 4 Share The final week of the semester coincides with the annual AI Winter megaconference, and so I’ll lean into posting about AI this week. And given the spiritual syzygy, I’m going to focus on a fascinating intersection between AI methods and AI culture : Reinforcement learning. Reinforcement learning is a weird subfield of AI that has plagued me since I got to Berkeley. I have written about it countless times. I have an old, pre-substack blog series about it. I wrote a whole sur
Explore this link on the map →saved by
related reading
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward is not the optimization target — LessWronglesswrong.com
- Part 1: Key Concepts in RL - Spinning Up documentationspinningup.openai.com
- Evolution as Backstop for Reinforcement Learning · Gwern.netgwern.net
- Debugging Reinforcement Learning Systemsandyljones.com
- Reinforcement learning - Wikipediaen.wikipedia.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- Just Ask for Generalization | Eric Jangevjang.com
- Deep Reinforcement Learning: Pong from Pixelskarpathy.github.io
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- RLHF Bookrlhfbook.com
- An Updated Introduction to Reinforcement Learning | Sri's Blogsrianumakonda.com