flâneur — a map of the web's best reading

blog.alexalemi.com Why KL?

blog.alexalemi.com · 5,388 words · saved by 1 readers

The Kullback-Liebler divergence, or KL divergence, or relative entropy, or relative information, or information gain, or expected weight of evidence, or information divergence (it goes by a lot of different names) is unique among the ways to measure the difference between two probability distributions. It holds a special and privileged place, being used to define all of the core concepts in information theory, such as mutual information. Why is the relative information so special and where does it come from? How should you interpret it? What is a nat anyway? In this note, I'll try to give a better understanding and set of intuitions about what KL is, why it's interesting, where it comes from and what it's good for. Let's see if we can motivate the form of the KL axiomatically. Imagine we have some prior set of beliefs summarized as a probability distribution 𝑞 . In light of some kind of evidence, we update our beliefs to a new distribution 𝑝 . How much did we update our beliefs? Ho

Why KL? Alexander A. Alemi. 2020-08-07 The Kullback-Liebler divergence , or KL divergence, or relative entropy, or relative information, or information gain, or expected weight of evidence, or information divergence (it goes by a lot of different names) is unique among the ways to measure the difference between two probability distributions. It holds a special and privileged place, being used to define all of the core concepts in information theory, such as mutual information. Why is the relative information so special and where does it come from? How should you interpret it? What is a nat any

Explore this link on the map →

related reading