✳flâneur — a map of the web's best reading
Small Models Can Introspect, Too
vgel.me · 7,794 words · saved by 1 readers
Doing introspection experiments with an OSS model.
Small Models Can Introspect, Too Posted December 12, 2025 Recent work by Anthropic showed that Claude models, primarily Opus 4 and Opus 4.1, are able to introspect--detecting when external concepts have been injected into their activations. But not all of us have Opus at home! By looking at the logits, we show that a 32B open-source model that at first appears unable to introspect actually is subtly introspecting. We then show that better prompting can significantly improve introspection performance, and throw the logit lens and emergent misalignment into the mix, showing that the model can in
Explore this link on the map →related reading
- Emergent Introspective Awareness in Large Language Modelstransformer-circuits.pub
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- Detecting the Disturbance: A Nuanced View of Introspective Abilities in LLMsarxiv.org
- Introspection or entropy? Re-examining concept-injection “introspection” in open models — LessWronglesswrong.com
- [2601.01828] Emergent Introspective Awareness in Large Language Modelsarxiv.org
- [2602.20031] Latent Introspection: Models Can Detect Prior Concept Injectionsarxiv.org
- [2601.01828] Emergent Introspective Awareness in Large Language Modelsarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org
- [2410.13787] Looking Inward: Language Models Can Learn About Themselves by Introspectionarxiv.org