✳flâneur — a map of the web's best reading
DeepSeek R1's recipe to replicate o1 and the future of reasoning LMs
interconnects.ai · 3,455 words · saved by 1 readers
Yes, ring the true o1 replication bells for DeepSeek R1 🔔🔔🔔. Where we go next.
DeepSeek R1's recipe to replicate o1 and the future of reasoning LMs Yes, ring the true o1 replication bells for DeepSeek R1 🔔🔔🔔. Where we go next. Nathan Lambert Jan 21, 2025 236 2 34 Share Article voiceover 0:00 -19:32 Audio playback is not supported on your browser. Please upgrade. I have a few shows to share with you this week: On The Retort a week or two ago, we discussed the nature of AI and if it is a science (in the Kuhn’ian sense) I appeared on Dean W. Ball and Timothy B. Lee ’s new podcast AI Summer to discuss “thinking models” and the border between post-training and reasoning me
Explore this link on the map →saved by
related reading
- DeepSeek-R1arxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Understanding Reasoning LLMs - by Sebastian Raschka, PhDsebastianraschka.com
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- o1 and Reasoning | AndoLogsblog.ando.ai
- [2501.12948] DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learningarxiv.org
- GitHub - deepseek-ai/DeepSeek-R1 · GitHubgithub.com
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org
- Explore | alphaXivalphaxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Why reasoning models will generalize - by Nathan Lambertinterconnects.ai