[2502.04327] Value-Based Deep RL Scales Predictably
Abstract:Scaling data and compute is critical to the success of modern ML. However, scaling demands predictability: we want methods to not only perform well with more compute or data, but also have their performance be predictable from small-scale runs, without running the large-scale experiment. In this paper, we show that value-based off-policy RL methods are predictable despite community lore regarding their pathological behavior. First, we show that data and compute requirements to attain a given performance level lie on a Pareto frontier, controlled by the updates-to-data (UTD) ratio. By estimating this frontier, we can predict this data requirement when given more compute, and this compute requirement when given more data. Second, we determine the optimal allocation of a total resource budget across data and compute for a given performance and use it to determine hyperparameters that maximize performance for a given budget. Third, this scaling is enabled by first estimating predictable relationships between hyperparameters, which is used to manage effects of overfitting and plasticity loss unique to RL. We validate our approach using three algorithms: SAC, BRO, and PQL on DeepMind Control, OpenAI gym, and IsaacGym, when extrapolating to higher levels of data, compute, budget, or performance.
View PDF HTML (experimental) Abstract:Scaling data and compute is critical to the success of modern ML. However, scaling demands predictability: we want methods to not only perform well with more compute or data, but also have their performance be predictable from small-scale runs, without running the large-scale experiment. In this paper, we show that value-based off-policy RL methods are predictable despite community lore regarding their pathological behavior. First, we show that data and compute requirements to attain a given performance level lie on a Pareto frontier, controlled by the…
saved by
related reading
- Scaling Laws for Value-Based RLvalue-scaling.github.io
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- Q-learning is not yet scalableseohong.me
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- The Scaling Hypothesis · Gwern.netgwern.net
- [2510.13786] The Art of Scaling Reinforcement Learning Compute for LLMsarxiv.org
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- [2602.05999] On the Role of Computation in Reinforcement Learningarxiv.org
- [2606.05555] Representation Learning Enables Scalable Multitask Deep Reinforcement Learningarxiv.org
- [1709.06560] Deep Reinforcement Learning that Mattersarxiv.org