Introspection or entropy? Re-examining concept-injection “introspection” in open models — LessWrong
lesswrong.com · 5,058 words · saved by 1 readers
Thanks to Joshua Joseph, Dillon Plunkett, and Julian Huang for their feedback and for helping me refine these ideas. …
x Introspection or entropy? Re-examining concept-injection “introspection” in open models — LessWrong Interpretability (ML & AI) Introspection Philosophy Transformers AI Frontpage 62 Introspection or entropy? Re-examining concept-injection “introspection” in open models by agastyasridharan 25th Jun 2026 17 min read 4 62 Thanks to Joshua Joseph, Dillon Plunkett, and Julian Huang for their feedback and for helping me refine these ideas. Anthropic recently reported that language models can “introspect.” They take a steering vector for a concept like “oceans,” add it into the model’s internal acti
related reading
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- Small Models Can Introspect, Toovgel.me
- [2602.20031] Latent Introspection: Models Can Detect Prior Concept Injectionsarxiv.org
- Emergent Introspective Awareness in Large Language Modelstransformer-circuits.pub
- Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMsarxiv.org
- Emergent Introspective Awareness in Large Language Modelstransformer-circuits.pub
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- A global workspace in language models \ Anthropicanthropic.com
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com
- Prompt Injection as Role Confusionrole-confusion.github.io