Understanding Optimization in Deep Learning with Central Flows
This is the companion website for the paper Understanding Optimization in Deep Learning with Central Flows, published at ICLR 2025. This work takes a step towards a theory of optimization in deep learning. Traditional theories of optimization cannot describe the dynamics of optimization in deep learning, even in the simple setting of deterministic (i.e. full-batch) training. The challenge is that optimizers typically operate in a complex oscillatory regime termed the edge of stability. In this work, we develop theory that can describe the dynamics of optimization in this regime. Our key insight is that while the exact dynamics of an oscillatory optimizer may be challenging to analyze, the time-averaged (i.e. locally smoothed) dynamics are often much easier to understand. We characterize these dynamics with a central flow: a differential equation that directly models the time-averaged trajectory of an oscillatory optimizer, as illustrated in the following cartoon. This site is organized