Extracting Training Data from ChatGPT
We have just released a paper that allows us to extract several megabytes of ChatGPT’s training data for about two hundred dollars. (Language models, like ChatGPT, are trained on data taken from the public internet. Our attack shows that, by querying the model, we can actually extract some of the exact data it was trained on.) We estimate that it would be possible to extract ~a gigabyte of ChatGPT’s training dataset from the model by spending more money querying the model.
Extracting Training Data from ChatGPT We have just released a paper that allows us to extract several megabytes of ChatGPT’s training data for about two hundred dollars. (Language models, like ChatGPT, are trained on data taken from the public internet. Our attack shows that, by querying the model, we can actually extract some of the exact data it was trained on.) We estimate that it would be possible to extract ~a gigabyte of ChatGPT’s training dataset from the model by spending more money querying the model. Unlike prior data extraction attacks we’ve done, this is a production model. The key
Explore this link on the map →saved by
related reading
- Training is not the same as chatting: ChatGPT and other LLMs don’t remember everything you saysimonwillison.net
- What Is ChatGPT Doing … and Why Does It Work?-Stephen Wolfram Writingswritings.stephenwolfram.com
- gpt-4.pdfcdn.openai.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- [2406.10209] Be like a Goldfish, Don’t Memorize! Mitigating Memorization in Generative LLMsar5iv.labs.arxiv.org
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Nicholas Carlininicholas.carlini.com
- arxiv.org/pdf/2505.24832arxiv.org
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoningarxiv.org