The Multi-Armed Bandit Problem and Its Solutions | Lil'Log
The algorithms are implemented for Bernoulli bandit in lilianweng/multi-armed-bandit. Exploitation vs Exploration The exploration vs exploitation dilemma exists in many aspects of our life. Say, your favorite restaurant is right around the corner. If you go there every day, you would be confident of what you will get, but miss the chances of discovering an even better option. If you try new places all the time, very likely you are gonna have to eat unpleasant food from time to time. Similarly, online advisors try to balance between the known most attractive ads and the new ads that might be even more successful.
Table of Contents Exploitation vs Exploration What is Multi-Armed Bandit? Definition Bandit Strategies ε-Greedy Algorithm Upper Confidence Bounds Hoeffding's Inequality UCB1 Bayesian UCB Thompson Sampling Case Study Summary References The algorithms are implemented for Bernoulli bandit in lilianweng/multi-armed-bandit . Exploitation vs Exploration # The exploration vs exploitation dilemma exists in many aspects of our life. Say, your favorite restaurant is right around the corner. If you go there every day, you would be confident of what you will get, but miss the chances of discovering an eve
Explore this link on the map →saved by
related reading
- The Bayes Banditfrancesco215.github.io
- Multi-armed bandit - Wikipediaen.wikipedia.org
- Epsilon-Greedy Algorithm in Reinforcement Learning - GeeksforGeeksgeeksforgeeks.org
- Contextual Bandits and the Exp4 Algorithm – Bandit Algorithmsbanditalgs.com
- Is Offline Decision Making Possible with Only Few Samples? Reliable Decisions in Data-Starved Bandits via Trust Region Enhancementarxiv.org
- An Overview of Contextual Bandits | Towards Data Sciencetowardsdatascience.com
- Bayesian Bandits - optimizing click throughs with statistics - Chris Stucchiochrisstucchio.com
- Corruption-tolerant bandit learning | Machine Learning | Springer Nature Linklink.springer.com
- Exploration Strategies in Deep Reinforcement Learning | Lil'Loglilianweng.github.io
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- Brian Christian on computer science algorithms that tackle fundamental and universal problems — and whether they can help us live better in practice | 80,000 Hours80000hours.org
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io