LLMs can learn about themselves by introspection — LessWrong
TLDR: We find that LLMs are capable of introspection on simple tasks. We discuss potential implications of introspection for interpretability and the moral status of AIs. Paper Authors: Felix Binder, James Chua, Tomek Korbak, Henry Sleight, John Hughes, Robert Long, Ethan Perez, Miles Turpin, Owain Evans This post contains edited extracts from the full paper. Humans acquire knowledge by observing the external world, but also by introspection. Introspection gives a person privileged access to their current state of mind (e.g., thoughts and feelings) that is not accessible to external observers. Can LLMs introspect? We define introspection as acquiring knowledge that is not contained in or derived from training data but instead originates from internal states. Such a capability could enhance model interpretability. Instead of painstakingly analyzing a model's internal workings, we could simply ask the model about its beliefs, world models, and goals. More speculatively, an introspect
x LLMs can learn about themselves by introspection — LessWrong AI Frontpage 110 LLMs can learn about themselves by introspection by Felix J Binder , Owain_Evans 18th Oct 2024 AI Alignment Forum 11 min read 38 110 Ω 40 Are LLMs capable of introspection, i.e. special access to their own inner states? Can they use this access to report facts about themselves that are not in the training data? Yes — in simple tasks at least! TLDR: We find that LLMs are capable of introspection on simple tasks. We discuss potential implications of introspection for interpretability and the moral status of AIs. Pape
Explore this link on the map →related reading
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2506.05068] Does It Make Sense to Speak of Introspection in Large Language Models?arxiv.org
- [2508.14802] Privileged Self-Access Matters for Introspection in AIarxiv.org
- Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMsarxiv.org
- Emergent Introspective Awareness in Large Language Modelstransformer-circuits.pub
- Small Models Can Introspect, Toovgel.me
- [2508.14802] Privileged Self-Access Matters for Introspection in AIarxiv.org
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com