Contextual Bandits and the Exp4 Algorithm – Bandit Algorithms
In most bandit problems there is likely to be some additional information available at the beginning of rounds and often this information can potentially help with the action choices. For example, …
In most bandit problems there is likely to be some additional information available at the beginning of rounds and often this information can potentially help with the action choices. For example, in a web article recommendation system, where the goal is to keep the visitors engaged with the website, contextual information about the visitor of the website, the time of day, information on what is trendy, etc., can likely improve the choice of the article to be put on the “front-page”. For example, a science-oriented article is more likely to grab the attention of a science geek, and
Explore this link on the map →related reading
- The Multi-Armed Bandit Problem and Its Solutions | Lil'Loglilianweng.github.io
- The Bayes Banditfrancesco215.github.io
- An Overview of Contextual Bandits | Towards Data Sciencetowardsdatascience.com
- Multi-armed bandit - Wikipediaen.wikipedia.org
- Corruption-tolerant bandit learning | Machine Learning | Springer Nature Linklink.springer.com
- Lecture 1: Introduction to Sequence Prediction | CS 8803 Sequence Predictionthejakeyboy.github.io
- Bayesian Bandits - optimizing click throughs with statistics - Chris Stucchiochrisstucchio.com
- ARC progress update: Competing with sampling — LessWronglesswrong.com
- Epsilon-Greedy Algorithm in Reinforcement Learning - GeeksforGeeksgeeksforgeeks.org
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc
- Papers I’ve read this week, Mixture of Experts editionfinbarrtimbers.substack.com
- Real-time machine learning: challenges and solutionshuyenchip.com