Jiayi-Pan/TinyZero: Minimal reproduction of DeepSeek R1-Zero
github.com · 351 words · saved by 1 readers
Minimal reproduction of DeepSeek R1-Zero
⚠️ Deprecation Notice: This repo is no longer actively maintained. For running RL experiments, please directly use the latest veRL library. For the archived original documentation, see OLD_README.md. TinyZero is a reproduction of DeepSeek R1 Zero in countdown and multiplication tasks. We built upon veRL. Through RL, the 3B base LM develops self-verification and search abilities all on its own. You can experience the Aha moment yourself for < $30. Twitter thread: https://x.com/jiayi_pirate/status/1882839370505621655 Full experiment log: https://wandb.ai/jiayipan/TinyZero 📢: We release…
saved by
related reading
- DeepSeek-R1arxiv.org
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- As Rocks May Think | Eric Jangevjang.com
- Explore | alphaXivalphaxiv.org
- GitHub - deepseek-ai/DeepSeek-R1 · GitHubgithub.com
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Modelsarxiv.org
- Open-R1: a fully open reproduction of DeepSeek-R1huggingface.co
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org
- The Illustrated DeepSeek-R1newsletter.languagemodels.co
- Tinkerthinkingmachines.ai
- State of RL for reasoning LLMs | A. Weersaweers.de
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com