Exponential families from a single KL identity
Exponential families encompass the distributions central to modern machine learning — softmax, Gaussians, and Boltzmann distributions — and underlie the theory of variational inference, entropy-regularized reinforcement learning, and RLHF. We isolate a simple identity for exponential families that expresses the KL difference KL ( 𝑞 ∥ 𝑝 𝜆 2 ) − KL ( 𝑞 ∥ 𝑝 𝜆 1 ) in terms of the log-partition function 𝐴 ( 𝜆 ) and the moment 𝜇 𝑞 . Remarkably, this identity together with the single fact that KL ≥ 0 (with equality iff 𝑝 = 𝑞 ) suffices, by direct substitution and rearrangement, to derive a cluster of results that are classically obtained by separate, heavier arguments: a generalized three-point identity for arbitrary reference distributions, Pythagorean theorems for I-projections and reverse I-projections, convexity of the log-partition function, identification of its Legendre dual in KL terms, the Gibbs variational principle, and the explicit optimizer in KL-regular
\affiliations ∗ Scientific Consultant \website marc.dymetman@gmail.com \websiteref Exponential families from a single KL identity Abstract Exponential families encompass the distributions central to modern machine learning — softmax, Gaussians, and Boltzmann distributions — and underlie the theory of variational inference, entropy-regularized reinforcement learning, and RLHF. We isolate a simple identity for exponential families that expresses the KL difference KL ( q ∥ p λ 2 ) − KL ( q ∥ p λ 1 ) \mathrm{KL}(q\|p_{\lambda_{2}})-\mathrm{KL}(q\|p_{\lambda_{1}}) in terms of the log-partition
Explore this link on the map →saved by
related reading
- Maximum Entropy Methods (MaxEnt)bactra.org
- Kullback–Leibler divergence - Wikipediaen.wikipedia.org
- Six (and a half) intuitions for KL divergence — LessWronglesswrong.com
- Exponential family - Wikipediaen.wikipedia.org
- Approximating KL Divergencejoschu.net
- Andy Jonesandrewcharlesjones.github.io
- RL with KL penalties is better seen as Bayesian inference — LessWronglesswrong.com
- Short Notes on Divergence Measuresdanilorezende.com
- Visual Information Theory -- colah's blogcolah.github.io
- blog.alexalemi.com Why KL?blog.alexalemi.com
- blog.alexalemi.com KL is All You Needblog.alexalemi.com
- Gregory Gundersengregorygundersen.com