Recovering the Pre-Fine-Tuning Weights of Generative Models
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. The dominant paradigm in generative modeling consists of two steps: i) pre-training on a large-scale but unsafe dataset, ii) aligning the pre-trained model with human values via fine-tuning. This practice is considered safe, as no current method can recover the unsafe, pre-fine-tuning model weights. In this paper, we demonstrate that this assumption is often false. Concretely
Recovering the Pre-Fine-Tuning Weights of Generative Models Eliahu Horwitz Jonathan Kahana Yedid Hoshen School of Computer Science and Engineering The Hebrew University of Jerusalem, Israel https://vision.huji.ac.il/spectral_detuning/ {eliahu.horwitz, jonathan.kahana, yedid.hoshen}@mail.huji.ac.il Abstract The dominant paradigm in generative modeling consists of two steps: i) pre-training on a large-scale but unsafe dataset, ii) aligning the pre-trained model with human values via fine-tuning. This practice is considered safe, as no current method can recover the unsafe, pre-fine-tuning model
Explore this link on the map →related reading
- LoRA Without Regret - Thinking Machines Labthinkingmachines.ai
- LoRA vs Full Fine-tuning: An Illusion of Equivalencearxiv.org
- [2106.09685] LoRA: Low-Rank Adaptation of Large Language Modelsarxiv.org
- Parameter-Efficient LLM Finetuning With Low-Rank Adaptation (LoRA) - Lightning AIlightning.ai
- Reinforcement Learning Finetunes Small Subnetworks in Large Language Modelsarxiv.org
- Anatomy of a Modern Finetuning APIbenanderson.work
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- Recent Advances in Language Model Fine-tuningruder.io
- GenAI Handbookgenai-handbook.github.io
- Modern Pretraining Strategies: A Hands-On Guidetheneuralmaze.substack.com
- In-depth guide to fine-tuning LLMs with LoRA and QLoRA | Mercity Researchmercity.ai
- Finetune LLMs on your own consumer hardware using tools from PyTorch and Hugging Face ecosystem – PyTorchpytorch.org