ali on X: "22580: From GPT2 to Kimi3, Explained" / X
x.com · 2,847 words · saved by 4 readers
https://t.co/aivc7B4A7Z
Twenty-two thousand five hundred and eighty. That’s how many GPT-2 (2019) models fit inside KimiK3 (2026). We scaled up by a factor of 22,580 in seven years. But is it just... scale? In this worklog, I’ll walk through how we got here and how much, or how little, has actually changed since then. We’ll trace the major architectural developments leading to KimiK3. GPT-2 GPT-2 is a decoder-only architecture: The input receives token and positional embeddings: Each transformer block, zoomed in, looks like this: The attention process: Once the final hidden-state matrix is produced, the…
saved by
related reading
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- Kimi Linear: An Expressive, Efficient Attention Architecturearxiv.org
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- [2604.18002] Neural Garbage Collection: Learning to Forget while Learning to Reasonarxiv.org
- Kimi Linear: An Expressive, Efficient Attention Architecturealphaxiv.org
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- How LLMs Actually Work | 0xkato0xkato.xyz
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- LLM Architecture Gallery | Sebastian Raschka, PhDsebastianraschka.com
- Transformer Inference Arithmetic | kipply's blogkipp.ly
- DeltaNet Explained (Part I) | Songlin Yangsustcsonglin.github.io
- [2510.26692] Kimi Linear: An Expressive, Efficient Attention Architecturearxiv.org