flâneur — a map of the web's best reading

Seohong Park on X: "We scaled up an "alternative" paradigm in RL: *divide and conquer*. Compared to Q-learning (TD learning), divide and conquer can naturally scale to much longer horizons. Blog post: https://t.co/xtXBzya0bI Paper: https://t.co/nqYkLucsWu ↓ https://t.co/XCdgUzaLxF" / X

x.com · 167 words · saved by 1 readers

To view keyboard shortcuts, press question mark View keyboard shortcuts Home Explore 1 Notifications Chat Grok Premium Bookmarks Creator Studio Articles Profile More Post sean lee @infinitefun_ Post See new posts Conversation Seohong Park @seohong_park We scaled up an "alternative" paradigm in RL: *divide and conquer*. Compared to Q-learning (TD learning), divide and conquer can naturally scale to much longer horizons. Blog post: https:// seohong.me/blog/rl-withou t-td-learning … Paper: https:// arxiv.org/abs/2510.22512 ↓ 12:49 PM · Oct 29, 2025 · 75.7K Views 11 90 504 358 Relevant View quotes Post your reply Reply Seohong Park @seohong_park · Oct 29, 2025 Our problem setting is off-policy RL. That is, we want to train an RL policy without always requiring fresh samples, unlike on-policy methods like PPO and GRPO. The most widely used off-policy RL method is Q-learning (≈ temporal difference (TD) learning). 1 1 16 2.8K Seohong Park @seohong_park · Oct 29, 2025 The problem is

@seohong_park: We scaled up an "alternative" paradigm in RL: *divide and conquer*. Compared to Q-learning (TD learning), divide and conquer can naturally scale to much longer horizons. Blog post: https:// seohong.me/blog/rl-withou t-td-learning … Paper: https:// arxiv.org/abs/2510.22512 ↓ @seohong_park: Our problem setting is off-policy RL. That is, we want to train an RL policy without always requiring fresh samples, unlike on-policy methods like PPO and GRPO. The most widely used off-policy RL method is Q-learning (≈ temporal difference (TD) learning). @seohong_park: The problem is

Explore this link on the map →

saved by

related reading