The Illustrated DeepSeek-R1 - by Jay Alammar
newsletter.languagemodels.co · 1,634 words · saved by 1 readers
A recipe for reasoning LLMs
DeepSeek-R1 is the latest resounding beat in the steady drumroll of AI progress. For the ML R&D community, it is a major release for reasons including: It is an open weights model with smaller, distilled versions and It shares and reflects upon a training method to reproduce a reasoning model like OpenAI O1. In this post, we’ll see how it was built. Translations: Chinese, Korean, Turkish (Feel free to translate the post to your language and send me the link to add here) Contents: Recap: How LLMs are trained DeepSeek-R1 Training Recipe 1- Long chains of reasoning SFT Data 2- An…
related reading
- DeepSeek-R1arxiv.org
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- Open-R1: a fully open reproduction of DeepSeek-R1huggingface.co
- Understanding Reasoning LLMs - by Sebastian Raschka, PhDsebastianraschka.com
- As Rocks May Think | Eric Jangevjang.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- DeepSeek Proves AI Comes for All Jobs - Even AI Jobssubstack.com
- [2501.12948] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learningarxiv.org
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Modelsarxiv.org
- GitHub - deepseek-ai/DeepSeek-R1 · GitHubgithub.com
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com