flâneur — a map of the web's best reading

Understanding Memorization via Loss Curvature

goodfire.ai · 2,175 words · saved by 4 readers

Our new paper proposes a method to identify and suppress memorized content in models. This explainer provides an overview of our work. Jack Merullo* Srihita Vatsavaya* Lucius Bushnaq* Owen Lewis* Post by Srihita Vatsavaya and Michael Byun November 6, 2025 Read on arXiv → Language models memorize substantial parts of their training data. For example, prompting Llama 3.1 70B with Chapter ONE:⏎THE BOY can generate the entirety of Harry Potter and the Sorceror's Stone near-verbatim with high probability. From a certain point of view, it's not that surprising that a system trained as a next-token predictor will end up memorizing sequences of tokens. But that's a lot of verbatim memorization! And many fundamental questions about memorization in models remain mysterious: how are these “memories” stored? Are they localizable in model weights? Do they share common structure? How important is memorization to model capabilities? As a topical example of the latter question, consider the hypothesis

Understanding Memorization via Loss Curvature Research Understanding Memorization via Loss Curvature Our new paper proposes a method to identify and suppress memorized content in models. This explainer provides an overview of our work. Authors Jack Merullo * Siri (Srihita) Vatsavaya * Lucius Bushnaq * Owen Lewis * *Goodfire Post by Siri Vatsavaya and Michael Byun Published November 6, 2025 Full Paper Read on arXiv → Language models memorize substantial parts of their training data. For example, prompting Llama 3.1 70B with Chapter ONE:⏎THE BOY can generate the entirety of Harry Potter and the

Explore this link on the map →

saved by

related reading