frontier model training methodologies | Alex Wa’s Blog
How do labs train a frontier, multi-billion parameter model? We look towards seven open-weight frontier models: Hugging Face’s SmolLM3, Prime Intellect’s Intellect 3, Nous Research’s Hermes 4, OpenAI’s gpt-oss-120b, Moonshot’s Kimi K2, DeepSeek’s DeepSeek-R1, and Arcee’s Trinity series. This blog is an attempt at distilling the techniques, motivations, and considerations used to train their models with an emphasis on training methodology over infrastructure.
Share on: How do labs train a frontier, multi-billion parameter model? We look towards seven open-weight frontier models: Hugging Face’s SmolLM3 , Prime Intellect’s Intellect 3 , Nous Research’s Hermes 4 , OpenAI’s gpt-oss-120b , Moonshot’s Kimi K2 , DeepSeek’s DeepSeek-R1 , and Arcee’s Trinity series . This blog is an attempt at distilling the techniques, motivations, and considerations used to train their models with an emphasis on training methodology over infrastructure. These notes are largely structured based on Hugging Face’s SmolLM3 report due to its extensiveness, and it is currently
Explore this link on the map →saved by
related reading
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- How LLMs Actually Work | 0xkato0xkato.xyz
- The Little Book of Deep Learningfleuret.org
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- Composer2.pdfcursor.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Frontier language models have become much smaller | Epoch AIepoch.ai
- microgptkarpathy.github.io