J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
arxiv.org · 7,267 words · saved by 1 readers
N/A
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning J1: I NCENTIVIZING T HINKING IN LLM- AS - A -J UDGE VIA R EINFORCEMENT L EARNING Chenxi Whitehouse Tianlu Wang Ping Yu Xian Li Jason Weston Ilia Kulikov Swarnadeep Saha FAIR at Meta chenxwh@meta.com, swarnadeep@meta.com…
related reading
- DeepSeek-R1arxiv.org
- How to Use LLM as a Judge (Without Getting Burned)manthanguptaa.in
- LLM-as-a-judge: a complete guide to using LLMs for evaluationsevidentlyai.com
- As Rocks May Think | Eric Jangevjang.com
- The bitter lesson of LLM evalsparsed.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domainsarxiv.org
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- 2401.10020.pdfarxiv.org
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org