flâneur — a map of the web's best reading

RL without TD learning

seohong.me · 1,791 words · saved by 1 readers

In this post, I'll introduce a reinforcement learning (RL) algorithm based on an "alternative" paradigm: divide and conquer. Unlike traditional methods, this algorithm is not based on temporal difference (TD) learning (which has scalability challenges), and scales well to long-horizon tasks. Our problem setting is off-policy RL. Let's briefly review what this means. There are two classes of algorithms in RL: on-policy RL and off-policy RL. On-policy RL means we can only use fresh data collected by the current policy. In other words, we have to throw away old data each time we update the policy. Algorithms like PPO and GRPO (and policy gradient methods in general) belong to this category.[1] Off-policy RL means we don't have this restriction: we can use any kind of data, including old experience, human demonstrations, Internet data, and so on. So off-policy RL is more general and flexible than on-policy RL (and of course harder!). Q-learning is the most well-known off-policy RL algorith

RL without TD learning RL without TD learning Seohong Park UC Berkeley October 2025 \( \definecolor{myblue}{RGB}{89, 139, 231} \definecolor{plgray}{RGB}{153, 153, 153} \) In this post, I'll introduce a reinforcement learning (RL) algorithm based on an "alternative" paradigm: divide and conquer . Unlike traditional methods, this algorithm is not based on temporal difference (TD) learning (which has scalability challenges ), and scales well to long-horizon tasks. We can do RL based on divide and conquer, instead of TD learning. Problem setting: off-policy RL Our problem setting is off-policy RL

Explore this link on the map →

saved by

related reading