flâneur — a map of the web's best reading

Exponential families from a single KL identity

arxiv.org · 7,210 words · saved by 1 readers

Exponential families encompass the distributions central to modern machine learning — softmax, Gaussians, and Boltzmann distributions — and underlie the theory of variational inference, entropy-regularized reinforcement learning, and RLHF. We isolate a simple identity for exponential families that expresses the KL difference KL ​ ( 𝑞 ∥ 𝑝 𝜆 2 ) − KL ​ ( 𝑞 ∥ 𝑝 𝜆 1 ) in terms of the log-partition function 𝐴 ​ ( 𝜆 ) and the moment 𝜇 𝑞 . Remarkably, this identity together with the single fact that KL ≥ 0 (with equality iff 𝑝 = 𝑞 ) suffices, by direct substitution and rearrangement, to derive a cluster of results that are classically obtained by separate, heavier arguments: a generalized three-point identity for arbitrary reference distributions, Pythagorean theorems for I-projections and reverse I-projections, convexity of the log-partition function, identification of its Legendre dual in KL terms, the Gibbs variational principle, and the explicit optimizer in KL-regular

\affiliations ∗ Scientific Consultant \website marc.dymetman@gmail.com \websiteref Exponential families from a single KL identity Abstract Exponential families encompass the distributions central to modern machine learning — softmax, Gaussians, and Boltzmann distributions — and underlie the theory of variational inference, entropy-regularized reinforcement learning, and RLHF. We isolate a simple identity for exponential families that expresses the KL difference KL ​ ( q ∥ p λ 2 ) − KL ​ ( q ∥ p λ 1 ) \mathrm{KL}(q\|p_{\lambda_{2}})-\mathrm{KL}(q\|p_{\lambda_{1}}) in terms of the log-partition

Explore this link on the map →

saved by

related reading