Reinforcement Learning Finetunes Small Subnetworks in Large Language Models
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.
Reinforcement Learning Finetunes Small Subnetworks in Large Language Models Sagnik Mukherjee Lifan Yuan Dilek Hakkani-Tür Hao Peng University of Illinois Urbana-Champaign {sagnikm3,lifan4,dilek,haopeng}@illinois.edu Abstract Reinforcement learning (RL) yields substantial improvements in large language models’ (LLMs) downstream task performance and alignment with human values. Surprisingly, such large gains result from updating only a small subnetwork comprising just 5%-30% of the parameters, with the rest effectively unchanged. We refer to this phenomenon as parameter update sparsity induced b
Explore this link on the map →saved by
related reading
- LoRA vs Full Fine-tuning: An Illusion of Equivalencearxiv.org
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- State of RL for reasoning LLMs | A. Weersaweers.de
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- LLM Resourcesforrestbicker.com
- [2106.09685] LoRA: Low-Rank Adaptation of Large Language Modelsarxiv.org
- How We Build Trillion Parameter Reasoning RL with 10% GPUsmacaron.im
- Anatomy of a Modern Finetuning APIbenanderson.work
- Orbit - Ultra-efficient RL Pipelinespherelab.ai
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com