2310.12921.pdf
arxiv.org · 8,300 words · saved by 1 readers
N/A
Published as a conference paper at ICLR 2024 V ISION -L ANGUAGE M ODELS ARE Z ERO -S HOT R EWARD M ODELS FOR R EINFORCEMENT L EARNING Juan Rocamonde† ‡ Victoriano Montesinos Elvis Nava FAR AI Vertebra ETH AI Center Ethan Perez∗ David…
saved by
related reading
- RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedbackarxiv.org
- 2307.12950.pdfarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF | John Lambertjohnwlambert.github.io
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- Language Models can Solve Computer Tasksarxiv.org
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learningalphaxiv.org