✳flâneur — a map of the web's best reading
Nemotron 3 Ultra: what distillation can't fix
maximelabonne.substack.com · 1,354 words · saved by 1 readers
Ten specialist teachers distilled into one open 550B model
Nemotron 3 Ultra: what distillation can't fix Ten specialist teachers distilled into one open 550B model Maxime Labonne Jun 08, 2026 19 3 2 Share On June 4, 2026, Nvidia closed out its Nemotron 3 lineup with Nemotron 3 Ultra , a 550B-parameter MoE with 55B active and fully open weights. It is the most intelligent open-weight model from a US lab, but it still trails the Chinese open frontier led by Kimi K2.6. Nvidia's pitch is not the highest benchmark scores, but the highest throughput on long-running agentic workloads . Let's walk through the recipe and poke at the inference numbers behind th
Explore this link on the map →saved by
related reading
- [2605.23857] Strong Teacher Not Needed? On Distillation in LLM Pretrainingarxiv.org
- [2604.13010] Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillationarxiv.org
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- nemotron-3-super-120b-a12b Model by NVIDIA | NVIDIA NIMbuild.nvidia.com
- The Short Case for Nvidia Stock | YouTube Transcript Optimizeryoutubetranscriptoptimizer.com
- Composer2.pdfcursor.com
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- Nitrobrew: Fast, Lossless Distillation for Free | Tildeblog.tilderesearch.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- AINews | AINewsnews.smol.ai
- Nemotron_Diffusion_Tech_Report_v1.pdfd1qx31qr3h6wln.cloudfront.net