aiesi_post-training_public.pdf
kawine.github.io · 3,477 words · saved by 1 readers
N/A
POST-TRAINING LLMs Kawin Ethayarajh Assistant Professor of Applied AI - University of Chicago, Booth AI AND ECONOMICS SUMMER INSTITUTE 2026 August 6–11, 2026 • Chicago 1 The model you use is not the model that was pretrained. 2 A pretrained model continues text. USER MODEL How do I estimate the causal effect of a How do I estimate the causal effect of a minimum-wage increase?…
saved by
related reading
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- On-Policy Distillation - Thinking Machines Labthinkingmachines.ai
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- PostTrainBenchposttrainbench.com
- rlhfbook.com/book.pdfrlhfbook.com
- [2606.07527] Post-training is (Massive) Supervised Learningarxiv.org
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Elicitation, the simplest way to understand post-traininginterconnects.ai
- From REINFORCE to Dr. GRPOlancelqf.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- [2606.12360] Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signalarxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org