[2406.10209] Be like a Goldfish, Don’t Memorize! Mitigating Memorization in Generative LLMs
Large language models can memorize and repeat their training data, causing privacy and copyright risks. To mitigate memorization, we introduce a subtle modification to the next-token training objective that we call the…
\pdfcolInitStack tcb@breakable Be like a Goldfish, Don’t Memorize! Mitigating Memorization in Generative LLMs Abhimanyu Hans 1 , Yuxin Wen 1 , Neel Jain 1 , John Kirchenbauer 1 Hamid Kazemi 1 , Prajwal Singhania 1 , Siddharth Singh 1 , Gowthami Somepalli 1 Jonas Geiping 2,3 , Abhinav Bhatele 1 , Tom Goldstein 1 1 University of Maryland, 2 ELLIS Institute Tübingen, 3 Max Planck Institute for Intelligent Systems, Tübingen AI Center Correspondence to ahans1@umd.edu . Codebase: https://github.com/ahans30/goldfish-loss . Abstract Large language models can memorize and repeat their training data, ca
Explore this link on the map →related reading
- Understanding Memorization via Loss Curvaturegoodfire.ai
- Rethinking LLM Memorization – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- arxiv.org/pdf/2505.24832arxiv.org
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- Training is not the same as chatting: ChatGPT and other LLMs don’t remember everything you saysimonwillison.net
- GenAI Handbookgenai-handbook.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- [2605.15156] MeMo: Memory as a Modelarxiv.org
- Extracting Training Data from ChatGPTnot-just-memorization.github.io
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- VaultGemma: The world's most capable differentially private LLMresearch.google