Jane Street Blog - Does batch size matter?
This post is aimed at readers who are already familiar with stochastic gradient descent (SGD) and terms like “batch size”. For an introduction to these ideas...
Jane Street Blog - Does batch size matter? Does batch size matter? Oct 31, 2017 | 12 min read Share on Facebook Share on Twitter Share on LinkedIn By: Chris Hardin This post is aimed at readers who are already familiar with stochastic gradient descent (SGD) and terms like “batch size”. For an introduction to these ideas, I recommend Goodfellow et al.’s Deep Learning , in particular the introduction and, for more about SGD, Chapter 8. The relevance of SGD is that it has made it feasible to work with much more complex models than was formerly possible. There is a lot of talk about batch size in
Explore this link on the map →related reading
- [2101.12176] On the Origin of Implicit Regularization in Stochastic Gradient Descentarxiv.org
- [2603.21191] On the Role of Batch Size in Stochastic Conditional Gradient Methodsarxiv.org
- Stochastic gradient descent - Wikipediaen.m.wikipedia.org
- Why Momentum Really Worksdistill.pub
- Notes on the Origin of Implicit Regularization in SGDinference.vc
- The Little Book of Deep Learningfleuret.org
- [1706.02677] Accurate, Large Minibatch SGD: Training ImageNet in 1 Hourarxiv-vanity.com
- Maybe I was too harsh on deep learning theory (three days ago) — LessWronglesswrong.com
- [2503.22478] Almost Bayesian: The Fractal Dynamics of Stochastic Gradient Descentarxiv.org
- Why Does SGD Love Flat Minima? Marginally Better blogrishit-dagli.github.io
- Reiner Pope – The math behind how LLMs are trained and serveddwarkesh.com
- [1609.04836] On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minimaarxiv.org