The Bayes Bandit
Mobile note: this post is meant to be read on a computer. Several figures are interactive and use wide layouts, hover states, sliders, and precise clicks, so they may not work as intended on a phone or small tablet. This is one of the most important questions of the entire field of Artificial Intelligence. The whole field of dataset curation will go away once a good mathematical framework of curiosity emerges. Hopefully, this blog post will be the first part of a series of increasingly complex environments where I explain/derive the mathematical foundations of curiosity. The 𝐾 K-armed bandit problem is perhaps one of the worst-named pieces of math. However, it has important applications in reinforcement learning. Let's see how it works. You walk into a room with 𝐾 K slot machines. Each one pays out a noisy reward drawn from its own Gaussian distribution — you don't know the means, you don't know the variances, and you have a finite number of pulls. Your job is to earn as much money
The Bayes Bandit What is Curiosity? Mathematically? Authors Affiliations Francesco Sacco None Date 2025 Code github.com/Francesco215/Bayes-bandit Mobile note: this post is meant to be read on a computer. Several figures are interactive and use wide layouts, hover states, sliders, and precise clicks, so they may not work as intended on a phone or small tablet. This is one of the most important questions of the entire field of Artificial Intelligence. The whole field of dataset curation will go away once a good mathematical framework of curiosity emerges. Hopefully, this blog post will be the fi
Explore this link on the map →saved by
related reading
- The Multi-Armed Bandit Problem and Its Solutions | Lil'Loglilianweng.github.io
- [2510.13651] What is the objective of reasoning with reinforcement learning?arxiv.org
- Bayesian Bandits - optimizing click throughs with statistics - Chris Stucchiochrisstucchio.com
- Multi-armed bandit - Wikipediaen.wikipedia.org
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Is Offline Decision Making Possible with Only Few Samples? Reliable Decisions in Data-Starved Bandits via Trust Region Enhancementarxiv.org
- Epsilon-Greedy Algorithm in Reinforcement Learning - GeeksforGeeksgeeksforgeeks.org
- Reward is not the optimization target — LessWronglesswrong.com
- Approximately Bayesian Reasoning: Knightian Uncertainty, Goodhart, and the Look-Elsewhere Effect — LessWronglesswrong.com
- Geometric Exploration, Arithmetic Exploitation — LessWronglesswrong.com
- Contextual Bandits and the Exp4 Algorithm – Bandit Algorithmsbanditalgs.com
- An Overview of Contextual Bandits | Towards Data Sciencetowardsdatascience.com