Mosaic ResNet Deep Dive
TL;DR: We recently released a set of recipes which can accelerate training of a ResNet-50 on ImageNet by up to 7x over standard baselines. In this report we take a deep dive into the technical details of our work and share the insights we gained about optimizing the efficiency of model training over a broad range of compute budgets.
Mosaic ResNet Deep Dive | Databricks Blog Skip to main content TL;DR: We recently released a set of recipes which can accelerate training of a ResNet-50 on ImageNet by up to 7x over standard baselines. In this report we take a deep dive into the technical details of our work and share the insights we gained about optimizing the efficiency of model training over a broad range of compute budgets. Introduction ResNets have established themselves as the go-to baseline and testbed for computer vision research (see ResNet Strikes Back or the PyTorch blog ). More efficient training recipes can save m
Explore this link on the map →saved by
related reading
- Efficiently Estimating Pareto Frontiers with Cyclic Learning Rate Schedules | Databricks Blogmosaicml.com
- ⏯️ Autoresume Training - Composerdocs.mosaicml.com
- ♻️ Auto Microbatching - Composerdocs.mosaicml.com
- [1709.05011] ImageNet Training in Minutesarxiv-vanity.com
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- Making Deep Learning go Brrrr From First Principleshorace.io
- [1905.11946] EfficientNet: Rethinking Model Scaling for Convolutional Neural Networksarxiv.org
- Composer2.pdfcursor.com
- arxiv.org/pdf/2512.24880#page=3.56arxiv.org
- The Little Book of Deep Learningfleuret.org
- Mosaic LLMs: GPT-3 quality formosaicml.com
- Latest | Epoch AIepochai.org