✳flâneur — a map of the web's best reading
The Big LLM Architecture Comparison
magazine.sebastianraschka.com · 14,321 words · saved by 1 readers
From DeepSeek-V3 to Kimi K2: A Look At Modern LLM Architecture Design
The Big LLM Architecture Comparison From DeepSeek V3 to GLM-5: A Look At Modern LLM Architecture Design Sebastian Raschka, PhD Jul 19, 2025 1,996 97 175 Share Last updated: Apr 2, 2026 (added Gemma 4 in section 23) It has been seven years since the original GPT architecture was developed. At first glance, looking back at GPT-2 (2019) and forward to DeepSeek V3 and Llama 4 (2024-2025), one might be surprised at how structurally similar these models still are. Sure, positional embeddings have evolved from absolute to rotational (RoPE), Multi-Head Attention has largely given way to Grouped-Query
Explore this link on the map →related reading
- LLM Architecture Gallery | Sebastian Raschka, PhDsebastianraschka.com
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- How LLMs Actually Work | 0xkato0xkato.xyz
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- 2502.11089arxiv.org
- Bits, FLOPS, and Watts: A Systems-Level Perspective of Scaling LLMs — Part 1 | by Asheesh Goja | Mediummedium.com
- Understanding Attention in LLMs | Bartosz Milewski's Programming Cafebartoszmilewski.com
- LLM Resourcesforrestbicker.com
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev