Reinforcement Learning for Knowledge Awareness | kalomaze's kalomazing blog
Intro If you're active at all when it comes to language model research these days, you almost certainly have an implicit mental model for what kinds of tr...
Reinforcement Learning for Knowledge Awareness – kalomaze's kalomazing blog Reinforcement Learning for Knowledge Awareness 07 May, 2026 Intro If you're active at all when it comes to language model research these days, you almost certainly have an implicit mental model for what kinds of training have been useful for capability improvement in the post-o1 era. The paradigmatic shortlist is essentially as follows: Pretraining , which targets diverse conditional prediction on natural webtext (or its synthetic derivatives). Mid-training , which targets concentrated, high quality data mixes, typical
saved by
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- RLHF & Post-Training Course by Nathan Lambertrlhfbook.com
- RLHF | John Lambertjohnwlambert.github.io
- DeepSeek-R1arxiv.org
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- rlhfbook.com/book.pdfrlhfbook.com
- RLHF Bookrlhfbook.com
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- RL's Razor: Why Online Reinforcement Learning Forgets Lessarxiv.org
- Introduction | RLHF and Post-Training Book by Nathan Lambertrlhfbook.com