Golden Gate Claude \ Anthropic
When we turn up the strength of the “Golden Gate Bridge” feature, Claude’s responses begin to focus on the Golden Gate Bridge. For a short time, we’re making this model available for everyone to interact with.
UPDATE: Golden Gate Claude was online for a 24-hour period as a research demo and is no longer available. If you'd like to find out more about our research on interpretability and the activation of features within Claude, please see this post or our full research paper. On Tuesday, we released a major new research paper on interpreting large language models, in which we began to map out the inner workings of our AI model, Claude 3 Sonnet. In the “mind” of Claude, we found millions of concepts that activate when the model reads relevant text or sees relevant images, which we call “features”.…
saved by
related reading
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- On the Biology of a Large Language Modeltransformer-circuits.pub
- A global workspace in language models \ Anthropicanthropic.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Claude Code Cheat Sheetcc.storyfox.cz
- AI Learning Resources & Guides from Anthropic \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- Mapping the mind of a large language model \ Anthropicanthropic.com
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com
- Best practices for Claude Code - Claude Code Docsanthropic.com