flâneur

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

arxiv.org · 5,984 words · saved by 1 readers

N/A

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains Anisha Gunjal Anthony Wang* Elaine Lau Vaskar Nath Bing Liu Sean Hendryx Scale AI anisha.gunjal@scale.com…

saved by

related reading