Emergent Introspective Awareness in Large Language Models
We investigate whether large language models can introspect on their internal states. It is difficult to answer this question through conversation alone, as genuine introspection cannot be distinguished from confabulations. Here, we address this challenge by injecting representations of known concepts into a model’s activations, and measuring the influence of these manipulations on the model’s self-reported states. We find that models can, in certain scenarios, notice the presence of injected concepts and accurately identify them. Models demonstrate some ability to recall prior internal representations and distinguish them from raw text inputs. Strikingly, we find that some models can use their ability to recall prior intentions in order to distinguish their own outputs from artificial prefills. In all these experiments, Claude Opus 4 and 4.1, the most capable models we tested, generally demonstrate the greatest introspective awareness; however, trends across models are complex and sen
Emergent Introspective Awareness in Large Language Models Transformer Circuits Thread Emergent Introspective Awareness in Large Language Models Author Jack Lindsey Affiliations Anthropic Published October 29th, 2025 Correspondence to jacklindsey@anthropic.com We investigate whether large language models are aware of their own internal states. It is difficult to answer this question through conversation alone, as genuine introspection cannot be distinguished from confabulations. Here, we address this challenge by injecting representations of known concepts into a model’s activations, and measur
Explore this link on the map →related reading
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- [2601.01828] Emergent Introspective Awareness in Large Language Modelsarxiv.org
- [2601.01828] Emergent Introspective Awareness in Large Language Modelsarxiv.org
- Small Models Can Introspect, Toovgel.me
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMsarxiv.org
- [2506.05068] Does It Make Sense to Speak of Introspection in Large Language Models?arxiv.org
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- A global workspace in language models \ Anthropicanthropic.com
- On the Biology of a Large Language Modeltransformer-circuits.pub
- [2602.20031] Latent Introspection: Models Can Detect Prior Concept Injectionsarxiv.org