flâneur — a map of the web's best reading

Reinforcement Learning Finetunes Small Subnetworks in Large Language Models

arxiv.org · 7,194 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions.

Reinforcement Learning Finetunes Small Subnetworks in Large Language Models Sagnik Mukherjee Lifan Yuan Dilek Hakkani-Tür Hao Peng University of Illinois Urbana-Champaign {sagnikm3,lifan4,dilek,haopeng}@illinois.edu Abstract Reinforcement learning (RL) yields substantial improvements in large language models’ (LLMs) downstream task performance and alignment with human values. Surprisingly, such large gains result from updating only a small subnetwork comprising just 5%-30% of the parameters, with the rest effectively unchanged. We refer to this phenomenon as parameter update sparsity induced b

Explore this link on the map →

saved by

related reading