Mapping LLM attractor states — LessWrong
lesswrong.com · 1,466 words · saved by 1 readers
I’d love low filter (1) feedback on the method, and (2) takes on which elements are worth putting more work into. …
x Mapping LLM attractor states — LessWrong AI Frontpage 18 Mapping LLM attractor states by Adam Bricknell 22nd Feb 2026 4 min read 9 18 I’d love low filter (1) feedback on the method, and (2) takes on which elements are worth putting more work into. I’ve favoured brevity at the expense of detail. AMA. The GitHub repo is here . The idea and why it could matter Inspired by the spiritual bliss attractor state in Claude Sonnet 4, I attempt to map attractor states for a given LLM, and see how stable they are. This write up summarises a simple approach which could be scaled up to fully map a given m
related reading
- models have some pretty funny attractor states — LessWronglesswrong.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- Prompt Injection as Role Confusionrole-confusion.github.io
- Large Language Model: world models or surface statistics?thegradient.pub
- 2025: The year in LLMssimonwillison.net
- Taking LLMs Seriously (As Language Models) — LessWronglesswrong.com