flâneur

The Illustrated DeepSeek-R1 - by Jay Alammar

newsletter.languagemodels.co · 1,634 words · saved by 1 readers

A recipe for reasoning LLMs

DeepSeek-R1 is the latest resounding beat in the steady drumroll of AI progress. For the ML R&D community, it is a major release for reasons including: It is an open weights model with smaller, distilled versions and It shares and reflects upon a training method to reproduce a reasoning model like OpenAI O1. In this post, we’ll see how it was built. Translations: Chinese, Korean, Turkish (Feel free to translate the post to your language and send me the link to add here) Contents: Recap: How LLMs are trained DeepSeek-R1 Training Recipe 1- Long chains of reasoning SFT Data 2- An…

related reading