Deep Double Descent (cross-posted on OpenAI blog)
By Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever This is a lightly edited and expanded version of the following post on the OpenAI blog about the followi…
By Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever This is a lightly edited and expanded version of the following post on the OpenAI blog about the following paper . While I usually don’t advertise my own papers on this blog, I thought this might be of interest to theorists, and a good follow up to my prior post . I promise not to make a habit out of it. –Boaz TL;DR: Our paper shows that double descent occurs in conventional modern deep learning settings: visual classification in the presence of label noise (CIFAR 10, CIFAR 100) and machine
Explore this link on the map →related reading
- Understanding “Deep Double Descent” — LessWronglesswrong.com
- Double descent in human learning · Chris Saidchris-said.io
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- The Little Book of Deep Learningfleuret.org
- The Scaling Hypothesis · Gwern.netgwern.net
- A Recipe for Training Neural Networkskarpathy.github.io
- 4 – The Overfitting Iceberg – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- [2503.02113] Deep Learning is Not So Mysterious or Differentarxiv.org
- The Decade of Deep Learning | Leo Gaobmk.sh
- Superposition, Memorization, and Double Descenttransformer-circuits.pub
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com