Efficiently Estimating Pareto Frontiers with Cyclic Learning Rate Schedules
Benchmarking the tradeoff between model accuracy and training time is computationally expensive. Cyclic learning rate schedules can construct a tradeoff curve in a single training run. These cyclic tradeoff curves can be used to evaluate the effects of algorithmic choices on network training efficiency.
Efficiently Estimating Pareto Frontiers with Cyclic Learning Rate Schedules | Databricks Blog Skip to main content Benchmarking the tradeoff between model accuracy and training time is computationally expensive. Cyclic learning rate schedules can construct a tradeoff curve in a single training run. These cyclic tradeoff curves can be used to evaluate the effects of algorithmic choices on network training efficiency. This work is has been posted as an arXiv preprint: Fast Benchmarking of Accuracy vs. Training Time with Cyclic Learning Rates . Problem Statement: Producing Tradeoff Curves Efficie
Explore this link on the map →saved by
related reading
- Mosaic ResNet Deep Dive | Databricks Blogmosaicml.com
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- [2310.07831] Optimal Linear Decay Learning Rate Schedules and Further Refinementsarxiv.org
- Aman's AI Journal • Primers • Ilya Sutskever's Top 30aman.ai
- Composer2.pdfcursor.com
- The Little Book of Deep Learningfleuret.org
- Neural network training makes beautiful fractals | Jascha’s blogsohl-dickstein.github.io
- Latest | Epoch AIepochai.org
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- [1709.05011] ImageNet Training in Minutesarxiv-vanity.com
- Marketplace: my first attempt at training without backprop on GPU efficiently – Fang-Pen's coding notefangpenlin.com
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io