Rethinking LLM Memorization – Machine Learning Blog | ML@CMU | Carnegie Mellon University
A central question in the discussion of large language models (LLMs) concerns the extent to which they memorize their training data versus how they generalize to new tasks and settings. Most practitioners seem to (at least informally) believe that LLMs do some degree of both: they clearly memorize parts of the training data—for example, they are often able to reproduce large portions of training data verbatim [Carlini et al., 2023]—but they also seem to learn from this data, allowing them to generalize to new settings. The precise extent to which they do one or the other has massive implications for the practical and legal aspects of such models [Cooper et al., 2023]. Do LLMs truly produce new content, or do they only remix their training data? Should the act of training on copyrighted data be deemed an unfair use of data, or should fair use be judged by some notion of model memorization? When dealing with humans, we distinguish plagiarizing content from learning from it, but how shoul
Rethinking LLM Memorization – Machine Learning Blog | ML@CMU | Carnegie Mellon University --> Input your search keywords and press Enter. Categories: Research Educational Categories: Research Educational machine learning Research Rethinking LLM Memorization Authors Avi Schwarzschild by Avi Schwarzschild --> Affiliations Published September 13, 2024 DOI Introduction A central question in the discussion of large language models (LLMs) concerns the extent to which they memorize their training data versus how they generalize to new tasks and settings. Most practitioners seem to (at least inf
Explore this link on the map →saved by
related reading
- arxiv.org/pdf/2505.24832arxiv.org
- Understanding Memorization via Loss Curvaturegoodfire.ai
- [2406.10209] Be like a Goldfish, Don’t Memorize! Mitigating Memorization in Generative LLMsar5iv.labs.arxiv.org
- LLM Daydreaming · Gwern.netgwern.net
- Training is not the same as chatting: ChatGPT and other LLMs don’t remember everything you saysimonwillison.net
- [2605.15156] MeMo: Memory as a Modelarxiv.org
- The Many Ways that Digital Minds Can Know – Ryan Moulton's Articlesmoultano.wordpress.com
- suchir.net/fair_use.htmlsuchir.net
- GenAI Handbookgenai-handbook.github.io
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- Reimagining LLM Memory: Using Context as Training Data Unlocks Models That Learn at Test-Time | NVIDIA Technical Blogdeveloper.nvidia.com
- [2501.00663] Titans: Learning to Memorize at Test Timearxiv.org