DeepSeek-V3: A Large-Scale MoE Pretraining Benchmark for MLPerf Training v6.0 - MLCommons
mlcommons.org · 736 words · saved by 1 readers
MLPerf Training v6.0 introduces a large-scale pretraining benchmark built on DeepSeek-V3, bringing Mixture-of-Experts (MoE) evaluation to the suite.
Motivation and Architectural Relevance As Large Language Model (LLM) development increasingly adopts sparse computation, the benchmarks used to evaluate training performance need to keep pace. MLPerf™ Training v6.0 adds a large-scale pretraining benchmark built on DeepSeek-V3, a Mixture-of-Experts (MoE) architecture with 671B total parameters, of which 37B are activated per token. This benchmark captures the performance of critical innovations now standard in the industry, including Multi-head Latent Attention (MLA) and auxiliary-loss–free load balancing. Technical Architecture &…
saved by
related reading
- Mixture of Experts Quantile Balancing: Validated at 32B-A5B (1e22 FLOPs) Scaleopenathena.ai
- [2607.16051] Loop the Loopies!arxiv.org
- [2509.14786] Pre-training under infinite computearxiv.org
- Mixture-of-Kittens: our open-source MoE megakernel for NVL72scursor.com
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Mixtral of Expertsarxiv.org
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- Composer2.pdfcursor.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Mixture of Experts Explainedhuggingface.co
- >10x More Efficient Pretraining — Magicmagic.dev