mazoure20a.pdf
proceedings.mlr.press · 6,837 words · saved by 1 readers
N/A
Leveraging exploration in off-policy algorithms via normalizing flows Bogdan Mazoure∗ 1,2 , Thang Doan* 1,2 , Audrey Durand2,3 , Joelle Pineau1,2,4 , R Devon Hjelm2,5,6 1 2 3 McGill University Mila – Quebec AI Institute Université Laval 4 5 Facebook AI Research Microsoft Research Montreal 6…
saved by
related reading
- [2506.22401] Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RLarxiv.org
- RL_Notes__final_.pdfjubayer-ibn-hamid.github.io
- [2602.11399] Can We Really Learn One Representation to Optimize All Rewards?arxiv.org
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org
- [1709.06560] Deep Reinforcement Learning that Mattersarxiv.org
- A Gallery of Methods Beyond RL — Part I: Sampling Methodsshengyu-feng.github.io
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- Soft Actor-Critic - Spinning Up documentationspinningup.openai.com
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Flow Matching Policy Gradientsflowreinforce.github.io
- Soft Actor-Critic for continuous and discrete actions | by Wah Loon Keng | Mediummedium.com