✳flâneur — a map of the web's best reading
SPAR-Self-Forecasting/gemma2-boolq-calibration: RL training of Gemma 2 2B IT for calibrated YES/NO probability estimates on BoolQ using GRPO ·
github.com · 576 words · saved by 1 readers
RL training of Gemma 2 2B IT for calibrated YES/NO probability estimates on BoolQ using GRPO
RL for Calibration: Gemma 2 2B on BoolQ Model : eruzak/gemma-2-2b-it-reasoning-high-boolq-calibration Train Gemma 2 2B IT to give calibrated YES/NO probability estimates on BoolQ questions using RL (GRPO via prime-rl ). The model reads a passage, reasons briefly, then outputs ANSWER: YES/NO with XX% probability . Reward = 1 − Brier score , with a bucket length penalty to encourage chain-of-thought reasoning without rambling. Results Training ran for 50 steps (~107 minutes) on 2× A100 80GB . Metric Step 0 Step 49 Reward ~0.38 0.911 ECE ~0.25 ~0.05 Accuracy ~0.70 ~0.86 Parse rate 100% 100% The m
Explore this link on the map →related reading
- State of RL for reasoning LLMs | A. Weersaweers.de
- DeepSeek-R1arxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Composer2.pdfcursor.com
- Tutorial: Train your own Reasoning model with GRPO | Unsloth Documentationdocs.unsloth.ai
- Understanding R1-Zero-Like Training: A Critical Perspectivearxiv.org
- Explore | alphaXivalphaxiv.org
- Gemma 2: Improving Open Language Models at a Practical Sizearxiv.org
- [2510.13651] What is the objective of reasoning with reinforcement learning?arxiv.org
- The State of Reinforcement Learning for LLM Reasoningsebastianraschka.com