[2403.06634] Stealing Part of a Production Language Model
Abstract:We introduce the first model-stealing attack that extracts precise, nontrivial information from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2. Specifically, our attack recovers the embedding projection layer (up to symmetries) of a transformer model, given typical API access. For under \$20 USD, our attack extracts the entire projection matrix of OpenAI's Ada and Babbage language models. We thereby confirm, for the first time, that these black-box models have a hidden dimension of 1024 and 2048, respectively. We also recover the exact hidden dimension size of the gpt-3.5-turbo model, and estimate it would cost under $2,000 in queries to recover the entire projection matrix. We conclude with potential defenses and mitigations, and discuss the implications of possible future work that could extend our attack.
Abstract:We introduce the first model-stealing attack that extracts precise, nontrivial information from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2. Specifically, our attack recovers the embedding projection layer (up to symmetries) of a transformer model, given typical API access. For under \$20 USD, our attack extracts the entire projection matrix of OpenAI's Ada and Babbage language models. We thereby confirm, for the first time, that these black-box models have a hidden dimension of 1024 and 2048, respectively. We also recover the exact hidden dimension s
Explore this link on the map →related reading
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Transformer Circuits Threadtransformer-circuits.pub
- Extracting Training Data from ChatGPTnot-just-memorization.github.io
- GPT in 60 Lines of NumPy | Jay Modyjaykmody.com
- GitHub - karpathy/nanochat: The best ChatGPT that $100 can buy. · GitHubgithub.com
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Large Language Model: world models or surface statistics?thegradient.pub
- llm-security/README.md at main · greshake/llm-security · GitHubgithub.com
- GitHub - SoyGema/pulling_ace · GitHubgithub.com
- [2310.18512] Preventing Language Models From Hiding Their Reasoningar5iv.labs.arxiv.org
- Red Teaming Language Models with Language Models — Google DeepMinddeepmind.com
- Actually, Othello-GPT Has A Linear Emergent World Representation - Neel Nandaneelnanda.io