Is Offline Decision Making Possible with Only Few Samples? Reliable Decisions in Data-Starved Bandits via Trust Region Enhancement
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. What can an agent learn in a stochastic Multi-Armed Bandit (MAB) problem from a dataset that contains just a single sample for each arm? Surprisingly, in this work, we demonstrate that even in such a data-starved setting it may still be possible to find a policy competitive with the optimal one. This paves the way to reliable decision-making in settings where critical decisions must be made by relying only on a handful of samples. Our analysis reveals that stochastic policies can be substantially better than deterministic ones for offline decision-making. Focusing on offline multi-armed bandits, we design an algorithm called Trust Region of Uncertainty for Stochastic policy enhancemenT (TRUST) which is quite different from the predominant value-based low
Is Offline Decision Making Possible with Only Few Samples? Reliable Decisions in Data-Starved Bandits via Trust Region Enhancement License: CC BY 4.0 arXiv:2402.15703v1 [cs.LG] 24 Feb 2024 Is Offline Decision Making Possible with Only Few Samples? Reliable Decisions in Data-Starved Bandits via Trust Region Enhancement Ruiqi Zhang Yuexiang Zhai Andrea Zanette Abstract What can an agent learn in a stochastic Multi-Armed Bandit (MAB) problem from a dataset that contains just a single sample for each arm? Surprisingly, in this work, we demonstrate that even in such a data-starved setting it may st
Explore this link on the map →related reading
- The Multi-Armed Bandit Problem and Its Solutions | Lil'Loglilianweng.github.io
- [2307.04354] Policy Finetuning in Reinforcement Learning via Design of Experiments using Offline Dataar5iv.labs.arxiv.org
- The Bayes Banditfrancesco215.github.io
- Exploration for the Efficient Deployment of Reinforcement Learning Agentsopenreview.net
- [2007.08202] Provably Good Batch Reinforcement Learning Without Great Explorationar5iv.labs.arxiv.org
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Reinforcement Learning in Newcomblike Problemsproceedings.neurips.cc
- [2204.05618] When Should We Prefer Offline Reinforcement Learning Over Behavioral Cloning?ar5iv.labs.arxiv.org
- Multi-armed bandit - Wikipediaen.wikipedia.org
- [2006.03647] Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimizationar5iv.labs.arxiv.org
- Just make the straw bigger | Joan Veljajoanvelja.com
- State of RL for reasoning LLMs | A. Weersaweers.de