Contextual Bandits and Reinforcement Learning | by Pavel Surmenok | Towards Data Science
towardsdatascience.com · 4,770 words · saved by 1 readers
If you develop personalization of user experience for your website or an app, contextual bandits can help you. Using contextual bandits…
An Overview of Contextual Bandits | Towards Data Science Skip to content An Overview of Contextual Bandits A dynamic approach to treatment personalization Ugur Yildirim Feb 2, 2024 23 min read Share Outline Introduction When To Use Contextual Bandits 2.1. Contextual Bandit vs Multi-Armed Bandit vs A/B Testing 2.2. Contextual Bandit vs Multiple MABs 2.3. Contextual Bandit vs Multi-Step Reinforcement Learning 2.4. Contextual Bandit vs Uplift Modeling Exploration and Exploitation in Contextual Bandits 3.1. ε -greedy 3.2. Upper Confidence Bound (UCB) 3.3. Thompson Sampling Contextual Bandit Algori
related reading
- The Multi-Armed Bandit Problem and Its Solutions | Lil'Loglilianweng.github.io
- course_stat_rl.pdfmit.edu
- The Bayes Banditfrancesco215.github.io
- Contextual Bandits and the Exp4 Algorithm – Bandit Algorithmsbanditalgs.com
- Real-time machine learning: challenges and solutionshuyenchip.com
- K V Subrahmanyamcmi.ac.in
- Bayesian Bandits - optimizing click throughs with statistics - Chris Stucchiochrisstucchio.com
- Speeding up RL with high-leverage samples | Applied Computeappliedcompute.com
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc
- SuttonBartoIPRLBook2ndEd.pdfweb.stanford.edu
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- A Gallery of Methods Beyond RL — Part I: Sampling Methodsshengyu-feng.github.io