[2602.20031] Latent Introspection: Models Can Detect Prior Concept Injections
Abstract:We uncover a latent capacity for introspection in a Qwen 32B model, demonstrating that the model can detect when concepts have been injected into its earlier context and identify which concept was injected. While the model denies injection in sampled outputs, logit lens analysis reveals clear detection signals in the residual stream, which are attenuated in the final layers. Furthermore, prompting the model with accurate information about AI introspection mechanisms can dramatically strengthen this effect: the sensitivity to injection increases massively (0.3% -> 39.9%) with only a 0.6% increase in false positives. Also, mutual information between nine injected and recovered concepts rises from 0.61 bits to 1.05 bits, ruling out generic noise explanations. Our results demonstrate models can have a surprising capacity for introspection and steering awareness that is easy to overlook, with consequences for latent reasoning and safety.
Abstract:We uncover a latent capacity for introspection in a Qwen 32B model, demonstrating that the model can detect when concepts have been injected into its earlier context and identify which concept was injected. While the model denies injection in sampled outputs, logit lens analysis reveals clear detection signals in the residual stream, which are attenuated in the final layers. Furthermore, prompting the model with accurate information about AI introspection mechanisms can dramatically strengthen this effect: the sensitivity to injection increases massively (0.3% -> 39.9%) with only a 0.
Explore this link on the map →related reading
- Small Models Can Introspect, Toovgel.me
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- Emergent Introspective Awareness in Large Language Modelstransformer-circuits.pub
- Introspection or entropy? Re-examining concept-injection “introspection” in open models — LessWronglesswrong.com
- [2601.01828] Emergent Introspective Awareness in Large Language Modelsarxiv.org
- Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMsarxiv.org
- [2601.01828] Emergent Introspective Awareness in Large Language Modelsarxiv.org
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- [2508.14802] Privileged Self-Access Matters for Introspection in AIarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- A global workspace in language models \ Anthropicanthropic.com
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub