[2508.14802] Privileged Self-Access Matters for Introspection in AI
Abstract:Whether AI models can introspect is an increasingly important practical question. But there is no consensus on how introspection is to be defined. Beginning from a recently proposed ''lightweight'' definition, we argue instead for a thicker one. According to our proposal, introspection in AI is any process which yields information about internal states through a process more reliable than one with equal or lower computational cost available to a third party. Using experiments where LLMs reason about their internal temperature parameters, we show they can appear to have lightweight introspection while failing to meaningfully introspect per our proposed definition.
Abstract:Whether AI models can introspect is an increasingly important practical question. But there is no consensus on how introspection is to be defined. Beginning from a recently proposed ''lightweight'' definition, we argue instead for a thicker one. According to our proposal, introspection in AI is any process which yields information about internal states through a process more reliable than one with equal or lower computational cost available to a third party. Using experiments where LLMs reason about their internal temperature parameters, we show they can appear to have lightweight int
Explore this link on the map →related reading
- [2508.14802] Privileged Self-Access Matters for Introspection in AIarxiv.org
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- LLMs can learn about themselves by introspection — LessWronglesswrong.com
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- [2602.20031] Latent Introspection: Models Can Detect Prior Concept Injectionsarxiv.org
- Small Models Can Introspect, Toovgel.me
- [2506.05068] Does It Make Sense to Speak of Introspection in Large Language Models?arxiv.org
- Emergent Introspective Awareness in Large Language Modelstransformer-circuits.pub
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2601.01828] Emergent Introspective Awareness in Large Language Modelsarxiv.org
- Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMsarxiv.org