✳flâneur — a map of the web's best reading
MAAC/attention_sac.py at master · shariqiqbal2810/MAAC
github.com · 2,247 words · saved by 1 readers
When others mention you, assign you, or request your review, GitHub will let them know that you have limited availability.
MAAC/algorithms/attention_sac.py at master · shariqiqbal2810/MAAC · GitHub Skip to content You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} shariqiqbal2810 / MAAC Public Notifications You must be signed in to change notification settings Fork 181 Star 807 Files Expand file tree master / attention_sac.py Copy path Blame More file actions Blame More file actions Latest commit History History H
Explore this link on the map →saved by
related reading
- Soft Actor-Critic - Spinning Up documentationspinningup.openai.com
- Policy Gradient Algorithms | Lil'Loglilianweng.github.io
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- A (Long) Peek into Reinforcement Learning | Lil'Loglilianweng.github.io
- Learning Beyond Gradientstrinkle23897.github.io
- Soft Actor-Critic for continuous and discrete actions | by Wah Loon Keng | Mediummedium.com
- Algorithms — Ray 2.55.1docs.ray.io
- A Graphic Guide to Implementing PPO for Atari Games | Towards Data Sciencetowardsdatascience.com
- GitHub - marlbenchmark/on-policy: This is the official implementation of Multi-Agent PPO (MAPPO). · GitHubgithub.com
- Macaron-V1-Preview: 749B MoL Agent Model post-trained from GLM5.1macaron.im
- Reward is not the optimization target — LessWronglesswrong.com
- [1707.06347] Proximal Policy Optimization Algorithmsarxiv.org