The Bayes Bandit
Mobile note: this post is meant to be read on a computer. Several figures are interactive and use wide layouts, hover states, sliders, and precise clicks, so they may not work as intended on a phone or small tablet. This is one of the most important questions of the entire field of Artificial Intelligence. The whole field of dataset curation will go away once a good mathematical framework of curiosity emerges. Hopefully, this blog post will be the first part of a series of increasingly complex environments where I explain/derive the mathematical foundations of curiosity. The 𝐾 K-armed bandit problem is perhaps one of the worst-named pieces of math. However, it has important applications in reinforcement learning. Let's see how it works. You walk into a room with 𝐾 K slot machines. Each one pays out a noisy reward drawn from its own Gaussian distribution — you don't know the means, you don't know the variances, and you have a finite number of pulls. Your job is to earn as much money
The Bayes Bandit What is Curiosity? Mathematically? Authors Affiliations Francesco Sacco None Date 2025 Code github.com/Francesco215/Bayes-bandit Mobile note: this post is meant to be read on a computer. Several figures are interactive and use wide layouts, hover states, sliders, and precise clicks, so they may not work as intended on a phone or small tablet. This is one of the most important questions of the entire field of Artificial Intelligence. The whole field of dataset curation will go away once a good mathematical framework of curiosity emerges. Hopefully, this blog post will be the fi
saved by
related reading
- The Multi-Armed Bandit Problem and Its Solutions | Lil'Loglilianweng.github.io
- Bayesian Bandits - optimizing click throughs with statistics - Chris Stucchiochrisstucchio.com
- Multi-armed bandit - Wikipediaen.wikipedia.org
- course_stat_rl.pdfmit.edu
- K V Subrahmanyamcmi.ac.in
- Knowing About Biases Can Hurt People — LessWronglesswrong.com
- Is Offline Decision Making Possible with Only Few Samples? Reliable Decisions in Data-Starved Bandits via Trust Region Enhancementarxiv.org
- A Gallery of Methods Beyond RL — Part I: Sampling Methodsshengyu-feng.github.io
- SuttonBartoIPRLBook2ndEd.pdfweb.stanford.edu
- What is AIXI?jan.leike.name
- Bayes-Optimal Strategiesemergentmind.com
- ARC progress update: Competing with sampling — LessWronglesswrong.com