GLM-5.3-Flash: Frontier Intelligence, Flash Cost
We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash incorporates several architectural improvements over GLM-5. For the first time, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. It also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Combined with our latest 30T-token multimodal pre-training corpus, these changes let GLM-5.3-Flash produce more intelligence with less compute. Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this tra
We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks. GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving…
saved by
related reading
- GLM-5.2: Built for Long-Horizon Tasksz.ai
- Composer2.pdfcursor.com
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- M*: A Modular, Extensible, Serving System for Multimodal Modelsai.stanford.edu
- M*: One Serving System for Any-to-Any Multimodal Modelsmstar.stanford.edu
- GLM-5.2 is the step change for open agentsinterconnects.ai
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Together AI | The AI Native Cloudtogether.ai
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- 2403.09611.pdfarxiv.org
- AINews | AINewsnews.smol.ai
- GLM-5.1: Towards Long-Horizon Tasksz.ai