flâneur

ali on X: "22580: From GPT2 to Kimi3, Explained" / X

x.com · 2,847 words · saved by 4 readers

https://t.co/aivc7B4A7Z

Twenty-two thousand five hundred and eighty. That’s how many GPT-2 (2019) models fit inside KimiK3 (2026). We scaled up by a factor of 22,580 in seven years. But is it just... scale? In this worklog, I’ll walk through how we got here and how much, or how little, has actually changed since then. We’ll trace the major architectural developments leading to KimiK3. GPT-2 GPT-2 is a decoder-only architecture: The input receives token and positional embeddings: Each transformer block, zoomed in, looks like this: The attention process: Once the final hidden-state matrix is produced, the…

saved by

related reading