Emergent introspective awareness in large language models \ Anthropic
anthropic.com · 3,615 words · saved by 5 readers
Research from Anthropic on the ability of large language models to introspect
Interpretability Signs of introspection in large language models Oct 29, 2025 Read the paper Have you ever asked an AI model what’s on its mind? Or to explain how it came up with its responses? Models will sometimes answer questions like these, but it’s hard to know what to make of their answers. Can AI systems really introspect—that is, can they consider their own thoughts? Or do they just make up plausible-sounding answers when they’re asked to do so? Understanding whether AI systems can truly introspect has important implications for their transparency and reliability. If models can accurat
saved by
related reading
- Emergent Introspective Awareness in Large Language Modelstransformer-circuits.pub
- Emergent Introspective Awareness in Large Language Modelstransformer-circuits.pub
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- [2601.01828] Emergent Introspective Awareness in Large Language Modelsarxiv.org
- [2601.01828] Emergent Introspective Awareness in Large Language Modelsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Small Models Can Introspect, Toovgel.me
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- A global workspace in language models \ Anthropicanthropic.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMsarxiv.org