✳flâneur — a map of the web's best reading
Scaling Laws for Value-Based RL
value-scaling.github.io · 4,597 words · saved by 1 readers
With the right design decisions, value-based RL admits predictable scaling.
Scaling Laws for Value-Based RL Compute-Optimal Scaling for Value-Based Deep RL August 2025 Preston Fu *, Oleh Rybkin *, Zhiyuan Zhou , Michal Nauman , Pieter Abbeel , Sergey Levine , Aviral Kumar NeurIPS , 2025 arXiv Code Thread Poster Value-Based Deep RL Scales Predictably February 2025 Oleh Rybkin , Michal Nauman , Preston Fu , Charlie Snell , Pieter Abbeel , Sergey Levine , Aviral Kumar ICML , 2025 ICLR Robot Learning Workshop , 2025 ( oral ) arXiv Code Thread Poster In the era of large-scale AI, it is important to prototype new training methodologies at small scales before running at larg
Explore this link on the map →saved by
related reading
- Q-learning is not yet scalableseohong.me
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- On neural scaling and the quanta hypothesisericjmichaud.com
- The Scaling Hypothesis · Gwern.netgwern.net
- [2510.13786] The Art of Scaling Reinforcement Learning Compute for LLMsarxiv.org
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- Just Ask for Generalization | Eric Jangevjang.com
- RL without TD learningseohong.me
- New Scaling Laws for Large Language Models — LessWronglesswrong.com
- [2406.09329] Is Value Learning Really the Main Bottleneck in Offline RL?ar5iv.labs.arxiv.org