Deep Double Descent (cross-posted on OpenAI blog)
By Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever This is a lightly edited and expanded version of the following post on the OpenAI blog about the followi…
By Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever This is a lightly edited and expanded version of the following post on the OpenAI blog about the following paper . While I usually don’t advertise my own papers on this blog, I thought this might be of interest to theorists, and a good follow up to my prior post . I promise not to make a habit out of it. –Boaz TL;DR: Our paper shows that double descent occurs in conventional modern deep learning settings: visual classification in the presence of label noise (CIFAR 10, CIFAR 100) and machine
related reading
- [2303.14151] Double Descent Demystified: Identifying, Interpreting & Ablating the Sources of a Deep Learning Puzzlearxiv.org
- Understanding “Deep Double Descent” — LessWronglesswrong.com
- Double descent in human learning · Chris Saidchris-said.io
- Scaling Laws, Carefully | Lil'Loglilianweng.github.io
- A Theory of Deep Learning | Elements of a Vector Spaceelonlit.com
- The Little Book of Deep Learningfleuret.org
- A Recipe for Training Neural Networkskarpathy.github.io
- The Scaling Hypothesis · Gwern.netgwern.net
- 4 – The Overfitting Iceberg – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- Understanding deep learning requires rethinking generalizationarxiv.org
- [2503.02113] Deep Learning is Not So Mysterious or Differentarxiv.org
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org