flâneur — a map of the web's best reading

frontier model training methodologies | Alex Wa’s Blog

djdumpling.github.io · 16,052 words · saved by 7 readers

How do labs train a frontier, multi-billion parameter model? We look towards seven open-weight frontier models: Hugging Face’s SmolLM3, Prime Intellect’s Intellect 3, Nous Research’s Hermes 4, OpenAI’s gpt-oss-120b, Moonshot’s Kimi K2, DeepSeek’s DeepSeek-R1, and Arcee’s Trinity series. This blog is an attempt at distilling the techniques, motivations, and considerations used to train their models with an emphasis on training methodology over infrastructure.

Share on: How do labs train a frontier, multi-billion parameter model? We look towards seven open-weight frontier models: Hugging Face’s SmolLM3 , Prime Intellect’s Intellect 3 , Nous Research’s Hermes 4 , OpenAI’s gpt-oss-120b , Moonshot’s Kimi K2 , DeepSeek’s DeepSeek-R1 , and Arcee’s Trinity series . This blog is an attempt at distilling the techniques, motivations, and considerations used to train their models with an emphasis on training methodology over infrastructure. These notes are largely structured based on Hugging Face’s SmolLM3 report due to its extensiveness, and it is currently

Explore this link on the map →

saved by

related reading