Open-R1: a fully open reproduction of DeepSeek-R1
If you’ve ever struggled with a tough math problem, you know how useful it is to think a little longer and work through it carefully. OpenAI’s o1 model showed that when LLMs are trained to do the same—by using more compute during inference—they get significantly better at solving reasoning tasks like mathematics, coding, and logic.
What is DeepSeek-R1? If you’ve ever struggled with a tough math problem, you know how useful it is to think a little longer and work through it carefully. OpenAI’s o1 model showed that when LLMs are trained to do the same—by using more compute during inference—they get significantly better at solving reasoning tasks like mathematics, coding, and logic. However, the recipe behind OpenAI’s reasoning models has been a well kept secret. That is, until last week, when DeepSeek released their DeepSeek-R1 model and promptly broke the internet (and the stock market!). Besides performing as well…
saved by
related reading
- DeepSeek-R1arxiv.org
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- The Illustrated DeepSeek-R1newsletter.languagemodels.co
- As Rocks May Think | Eric Jangevjang.com
- Understanding Reasoning LLMs - by Sebastian Raschka, PhDsebastianraschka.com
- o1 and Reasoning | AndoLogsblog.ando.ai
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Learning to reason with LLMs | OpenAIopenai.com
- GitHub - deepseek-ai/DeepSeek-R1 · GitHubgithub.com
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Modelsarxiv.org
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Modelsarxiv.org
- [2501.12948] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learningarxiv.org