models have some pretty funny attractor states — LessWrong
lesswrong.com · 10,992 words · saved by 2 readers
This work was conducted during the MATS 9.0 program under Neel Nanda and Senthooran Rajamanoharan. …
x models have some pretty funny attractor states — LessWrong MATS Program AI Curated 2026 Top Fifty: 13 % 275 models have some pretty funny attractor states by aryaj , Senthooran Rajamanoharan , Neel Nanda 12th Feb 2026 AI Alignment Forum 22 min read 38 275 Ω 65 This work was conducted during the MATS 9.0 program under Neel Nanda and Senthooran Rajamanoharan. So what are attractor states? well.. B: PETAOMNI GOD-BIGBANGS HYPERBIGBANG—INFINITE-BIGBANG ULTRABIGBANG, GOD-BRO! Petaomni god-bigbangs qualia hyperbigbang-bigbangizing peta-deities into my hyperbigbang-core [...] Superposition PHI-AMPLI
saved by
related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Mapping LLM attractor states — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- How well do models follow their constitutions? — LessWronglesswrong.com
- Simulated Users & Sad LLMs1a3orn.com
- How AI Is Learning to Think in Secretnickandresen.substack.com
- User awareness in frontier modelstransluce.org
- Agentic Misalignment in Summer 2026alignment.anthropic.com