✳flâneur — a map of the web's best reading
Emergent introspective awareness in large language models \ Anthropic
anthropic.com · 3,615 words · saved by 1 readers
Research from Anthropic on the ability of large language models to introspect
Interpretability Signs of introspection in large language models Oct 29, 2025 Read the paper Have you ever asked an AI model what’s on its mind? Or to explain how it came up with its responses? Models will sometimes answer questions like these, but it’s hard to know what to make of their answers. Can AI systems really introspect—that is, can they consider their own thoughts? Or do they just make up plausible-sounding answers when they’re asked to do so? Understanding whether AI systems can truly introspect has important implications for their transparency and reliability. If models can accurat
Explore this link on the map →related reading
- Emergent Introspective Awareness in Large Language Modelstransformer-circuits.pub
- [2601.01828] Emergent Introspective Awareness in Large Language Modelsarxiv.org
- [2601.01828] Emergent Introspective Awareness in Large Language Modelsarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Small Models Can Introspect, Toovgel.me
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMsarxiv.org
- A global workspace in language models \ Anthropicanthropic.com
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- [2506.05068] Does It Make Sense to Speak of Introspection in Large Language Models?arxiv.org
- LLMs can learn about themselves by introspection — LessWronglesswrong.com
- [2602.20031] Latent Introspection: Models Can Detect Prior Concept Injectionsarxiv.org