Mapping the Mind of a Large Language Model \ Anthropic
We have identified how millions of concepts are represented inside Claude Sonnet, one of our deployed large language models. This is the first ever detailed look inside a modern, production-grade large language model.
Interpretability Mapping the mind of a large language model May 21, 2024 Read the paper Today we report a significant advance in understanding the inner workings of AI models. We have identified how millions of concepts are represented inside Claude Sonnet, one of our deployed large language models. This is the first ever detailed look inside a modern, production-grade large language model. This interpretability discovery could, in future, help us make AI models safer. We mostly treat AI models as a black box: something goes in and a response comes out, and it's not clear why the model gave th
Explore this link on the map →related reading
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- A global workspace in language models \ Anthropicanthropic.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Tracing the Thoughts of a Large Language Model — LessWronglesswrong.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub