DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL
DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL Get Notion free DeepScaleR: Surpassing O1-Preview with a 1.5B Model by Scaling RL Michael Luo*, Sijun Tan*, Justin Wong†, Xiaoxiang Shi, William Tang, Manan Roongta, Colin Cai, Jeffrey Luo Advisors: Li Erran Li, Raluca Ada Popa, Ion Stoica *: Project Leads; †: Significant Contributor ✨ TL;DR RL magic is in the air! We introduce DeepScaleR-1.5B-Preview, a language model finetuned from Deepseek-R1-Distilled-Qwen-1.5B using simple reinforcement learning (RL). It achieves an impressive 43.1% Pass@1 accuracy on AIME2024 (+14.3% improvement over the base model), surpassing the performance of OpenAI’s o1-preview with just 1.5B parameters. We open sourced our dataset, code and training logs for everyone to progress on scaling intelligence with RL. 🌐 Website, 👨💻 Github, 🤗 HF Model, 🤗 HF Dataset, 📈 Wandb Logs, 🔎 Eval Logs DeepScaleR-1.5B-Preview Model AIME 2024 MATH 500 AMC 2023 Minerva Math Olympiad Bench Avg.
Notion’s local storage may be damaged. See (?) > Help & documentation > Reset Notion. Contact support if that doesn’t fix the issue. Your Firefox profile may be damaged. Visit https://firefox-storage-test.glitch.me/ to diagnose. Contact support if that doesn’t fix the issue. Your Chrome profile may be damaged. If you changed any chrome://flags then please reset them, then restart your browser. If issues persist, try making a new Chrome user. Contact support if that doesn’t fix the issue. Your Chrome profile may be damaged. For a more consistent experience download the Notion desktop app: https
Explore this link on the map →