flâneur — a map of the web's best reading

Post-training 101 | Tokens for Thoughts

tokens-for-thoughts.notion.site · 279 words · saved by 4 readers

Skip to content Post-training 101 Get Notion free Post-training 101 A hitchhiker's guide into LLM post-training EP1 of Tokens for Thoughts By Han Fang, Karthik Abinav Sankararaman Overview 1. The Journey from Pre‑training to an Instruct‑tuned Model 2. E2E Life Cycle of Post-training 3. What Is Supervised Fine-tuning? SFT Dataset 🔍 DATA EXAMPLES Data Quality in SFT Datasets How SFT Data Is Batched and Padded SFT Loss Function - Negative Log Likelihood Function Numerical Stability 4. What Are the Common RL Training Techniques? RL Rewards Reward Models and Preferences What does preference data look like? 🔍 DATA EXAMPLES Training rewards models RL Prompts and Data Verifiable rewards 🔍 DATA EXAMPLES Preference reward 🔍 DATA EXAMPLES Rubrics-guided rewards 🔍 DATA EXAMPLES RL Algorithms PPO GRPO 5. How Do We Evaluate Post-trained Models? Auto evaluation Ground truth based eval 🔍 DATA EXAMPLES LLM judge based eval 🔍 DATA EXAMPLES Human evaluation Point-wise eval 🔍 DATA EXAMPLES Prefere

Explore this link on the map →

saved by