frontier model training methodologies | Alex Wa’s Blog
How do labs train a frontier, multi-billion parameter model? We look towards seven open-weight frontier models: Hugging Face’s SmolLM3, Prime Intellect’s Intellect 3, Nous Research’s Hermes 4, OpenAI’s gpt-oss-120b, Moonshot’s Kimi K2, DeepSeek’s DeepSeek-R1, and Arcee’s Trinity series. This blog is an attempt at distilling the techniques, motivations, and considerations used to train their models with an emphasis on training methodology over infrastructure.
Share on: How do labs train a frontier, multi-billion parameter model? We look towards seven open-weight frontier models: Hugging Face’s SmolLM3 , Prime Intellect’s Intellect 3 , Nous Research’s Hermes 4 , OpenAI’s gpt-oss-120b , Moonshot’s Kimi K2 , DeepSeek’s DeepSeek-R1 , and Arcee’s Trinity series . This blog is an attempt at distilling the techniques, motivations, and considerations used to train their models with an emphasis on training methodology over infrastructure. These notes are largely structured based on Hugging Face’s SmolLM3 report due to its extensiveness, and it is currently
saved by
- Karan MJ
- Freeman Jiang
- Rishi Kothari
- Benedict Neo
- Akira Yoshiyama
- Denys
- Vihan Tiwari
- Ishaan Panigrahi
- Sam Wang
related reading
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- Composer2.pdfcursor.com
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- The Little Book of Deep Learningfleuret.org
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- laguna-m1-xs2-technical-report.pdfpoolside.ai
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- microgptkarpathy.github.io
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai