flâneur — a map of the web's best reading

Rethinking LLM Memorization – Machine Learning Blog | ML@CMU | Carnegie Mellon University

blog.ml.cmu.edu · 2,666 words · saved by 1 readers

A central question in the discussion of large language models (LLMs) concerns the extent to which they memorize their training data versus how they generalize to new tasks and settings. Most practitioners seem to (at least informally) believe that LLMs do some degree of both: they clearly memorize parts of the training data—for example, they are often able to reproduce large portions of training data verbatim [Carlini et al., 2023]—but they also seem to learn from this data, allowing them to generalize to new settings. The precise extent to which they do one or the other has massive implications for the practical and legal aspects of such models [Cooper et al., 2023]. Do LLMs truly produce new content, or do they only remix their training data? Should the act of training on copyrighted data be deemed an unfair use of data, or should fair use be judged by some notion of model memorization? When dealing with humans, we distinguish plagiarizing content from learning from it, but how shoul

Rethinking LLM Memorization – Machine Learning Blog | ML@CMU | Carnegie Mellon University --> Input your search keywords and press Enter. Categories: Research Educational Categories: Research Educational machine learning Research Rethinking LLM Memorization Authors Avi Schwarzschild by Avi Schwarzschild --> Affiliations Published September 13, 2024 DOI Introduction A central question in the discussion of large language models (LLMs) concerns the extent to which they memorize their training data versus how they generalize to new tasks and settings. Most practitioners seem to (at least inf

Explore this link on the map →

saved by

related reading