flâneur — a map of the web's best reading

On-Policy Distillation - Thinking Machines Lab

thinkingmachines.ai · 6,007 words · saved by 26 readers

On-policy, dense supervision is a useful tool for distillation

LLMs are capable of expert performance in focused domains, a result of several capabilities stacked together: perception of input, knowledge retrieval, plan selection, and reliable execution. This requires a stack of training approaches, which we can divide into three broad stages: Pre-training teaches general capacities such as language use, broad reasoning, and world knowledge. Mid-training imparts domain knowledge, such as code, medical databases, or internal company documents. Post-training elicits targeted behavior, such as instruction following, reasoning through math problems, or chat.

Explore this link on the map →

saved by

related reading