The Big LLM Architecture Comparison
magazine.sebastianraschka.com · 14,321 words · saved by 6 readers
From DeepSeek-V3 to Kimi K2: A Look At Modern LLM Architecture Design
The Big LLM Architecture Comparison From DeepSeek V3 to GLM-5: A Look At Modern LLM Architecture Design Sebastian Raschka, PhD Jul 19, 2025 1,996 97 175 Share Last updated: Apr 2, 2026 (added Gemma 4 in section 23) It has been seven years since the original GPT architecture was developed. At first glance, looking back at GPT-2 (2019) and forward to DeepSeek V3 and Llama 4 (2024-2025), one might be surprised at how structurally similar these models still are. Sure, positional embeddings have evolved from absolute to rotational (RoPE), Multi-Head Attention has largely given way to Grouped-Query
saved by
related reading
- ali (@waterloo_intern) on Xx.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- LLM Architecture Gallery | Sebastian Raschka, PhDsebastianraschka.com
- How LLMs Actually Work | 0xkato0xkato.xyz
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- LLM Resourcesforrestbicker.com
- Kimi Linear: An Expressive, Efficient Attention Architecturearxiv.org
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- DeepSeek-V3: A Large-Scale MoE Pretraining Benchmark for MLPerf Training v6.0mlcommons.org
- 2502.11089arxiv.org
- Paged Attention from First Principles: A View Inside vLLM – Hamza's Bloghamzaelshafie.bearblog.dev