Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
arxiv.org · 5,984 words · saved by 1 readers
N/A
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Anisha Gunjal Anthony Wang* Elaine Lau Vaskar Nath Bing Liu Sean Hendryx Scale AI anisha.gunjal@scale.com…
saved by
related reading
- [2605.12474] Reward Hacking in Rubric-Based Reinforcement Learningarxiv.org
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Reasoning in General Domains without Verifiersarxiv.org
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- Reinforcement Learning With Verifiable Rewards: How Data and Verifiers Shape RLVRsnorkel.ai
- Inverse Rubric Optimization: A testbed for agent science | Fulcrumfulcrum.inc
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- RLHF Bookrlhfbook.com
- LLM-as-a-Verifier: A General-Purpose Verification Framework | alphaXivalphaxiv.org
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com