flâneur — a map of the web's best reading

[2406.10209] Be like a Goldfish, Don’t Memorize! Mitigating Memorization in Generative LLMs

ar5iv.labs.arxiv.org · 8,040 words · saved by 1 readers

Large language models can memorize and repeat their training data, causing privacy and copyright risks. To mitigate memorization, we introduce a subtle modification to the next-token training objective that we call the…

\pdfcolInitStack tcb@breakable Be like a Goldfish, Don’t Memorize! Mitigating Memorization in Generative LLMs Abhimanyu Hans 1 , Yuxin Wen 1 , Neel Jain 1 , John Kirchenbauer 1 Hamid Kazemi 1 , Prajwal Singhania 1 , Siddharth Singh 1 , Gowthami Somepalli 1 Jonas Geiping 2,3 , Abhinav Bhatele 1 , Tom Goldstein 1 1 University of Maryland, 2 ELLIS Institute Tübingen, 3 Max Planck Institute for Intelligent Systems, Tübingen AI Center Correspondence to ahans1@umd.edu . Codebase: https://github.com/ahans30/goldfish-loss . Abstract Large language models can memorize and repeat their training data, ca

Explore this link on the map →

related reading