Reinforcement Learning as a fine-tuning paradigm | Ankesh Anand
Reinforcement Learning should be better seen as a “fine-tuning” paradigm that can add capabilities to general-purpose foundation models, rather than a paradigm that can bootstrap intelligence from scratch.
Reinforcement Learning (RL) should be better seen as a “fine-tuning” paradigm that can add capabilities to general-purpose pretrained models, rather than a paradigm that can bootstrap intelligence from scratch. Most contemporary reinforcement learning works involve training agents “tabula-rasa”, without relying on any sort of knowledge about the world. So when solving a task(s), the agent not only has to optimize the reward function at hand, but in the process also discover how to see, how physics works, what consequences its actions have, how language works, and so forth. This tends to work o
Explore this link on the map →related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- The Era of Experience Paper.pdfstorage.googleapis.com
- Just Ask for Generalization | Eric Jangevjang.com
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- rlhfbook.com/book.pdfrlhfbook.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Precise Manipulation with Efficient Online RLpi.website
- Reinforcement Learning for Knowledge Awareness – kalomaze's kalomazing blogkalomaze.bearblog.dev
- Reinforcement learning, AI, and general intelligenceartfintel.com
- [2603.21972] Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipearxiv.org