Scaling Laws for Value-Based RL
value-scaling.github.io · 4,597 words · saved by 1 readers
With the right design decisions, value-based RL admits predictable scaling.
Scaling Laws for Value-Based RL Compute-Optimal Scaling for Value-Based Deep RL August 2025 Preston Fu *, Oleh Rybkin *, Zhiyuan Zhou , Michal Nauman , Pieter Abbeel , Sergey Levine , Aviral Kumar NeurIPS , 2025 arXiv Code Thread Poster Value-Based Deep RL Scales Predictably February 2025 Oleh Rybkin , Michal Nauman , Preston Fu , Charlie Snell , Pieter Abbeel , Sergey Levine , Aviral Kumar ICML , 2025 ICLR Robot Learning Workshop , 2025 ( oral ) arXiv Code Thread Poster In the era of large-scale AI, it is important to prototype new training methodologies at small scales before running at larg
saved by
related reading
- Value-Based Deep RL Scales Predictablyarxiv.org
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Q-learning is not yet scalableseohong.me
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- [2510.13786] The Art of Scaling Reinforcement Learning Compute for LLMsarxiv.org
- The Scaling Hypothesis · Gwern.netgwern.net
- On neural scaling and the quanta hypothesisericjmichaud.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- Scaling is subtler than it seemsberen.io
- Just Ask for Generalization | Eric Jangevjang.com