flâneur — a map of the web's best reading

Introspection or entropy? Re-examining concept-injection “introspection” in open models — LessWrong

lesswrong.com · 5,058 words · saved by 1 readers

Thanks to Joshua Joseph, Dillon Plunkett, and Julian Huang for their feedback and for helping me refine these ideas. …

x Introspection or entropy? Re-examining concept-injection “introspection” in open models — LessWrong Interpretability (ML & AI) Introspection Philosophy Transformers AI Frontpage 62 Introspection or entropy? Re-examining concept-injection “introspection” in open models by agastyasridharan 25th Jun 2026 17 min read 4 62 Thanks to Joshua Joseph, Dillon Plunkett, and Julian Huang for their feedback and for helping me refine these ideas. Anthropic recently reported that language models can “introspect.” They take a steering vector for a concept like “oceans,” add it into the model’s internal acti

Explore this link on the map →

related reading