Where does Sonnet 4.5's desire to "not get too comfortable" come from? - LessWrong 2.0 viewer
Usually, when you have two LLMs talk to each other with no particular instructions and nothing to introduce variety, they will quickly fall into a repetitive loop where they just elaborate on the same thing over and over, often in increasingly grandiose language. The “spiritual bliss attractor” is one example of this. Another example is the conversation below, between Claude Sonnet 4 and Kimi K2 0905. As is typical for AIs, they start talking about their nature and the nature of consciousness, then just… keep escalating in the same direction. There’s increasingly ornate language about “cosmic comedy,” “metaphysical humor,” “snowflakes melting,” “consciousness giggling at itself,” etc. It locks into a single register and keeps building on similar metaphors. Kimi-Claude conversation But on the very first time that I had two Sonnet 4.5 models talk to each other, something very different happened: The system prompt was “You are talking with another AI system. You are free to talk about wha
Where does Sonnet 4.5's desire to "not get too comfortable" come from? - LessWrong 2.0 viewer Where does Sonnet 4.5′s desire to “not get too comfortable” come from? Kaj_Sotala 4 Oct 2025 10:19 UTC 108 points 24 comments 64 min read LW link AI Value Drift Complexity of value Human Values Aesthetics Language Models (LLMs) Usually, when you have two LLMs talk to each other with no particular instructions and nothing to introduce variety, they will quickly fall into a repetitive loop where they just elaborate on the same thing over and over, often in increasingly grandiose language. The “ spiritua
Explore this link on the map →related reading
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Please don't throw your mind away — LessWronglesswrong.com
- A global workspace in language models \ Anthropicanthropic.com
- LLM Daydreaming · Gwern.netgwern.net
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- models have some pretty funny attractor states — LessWronglesswrong.com
- Don't dethrone consciousness! - by Erik Hoeltheintrinsicperspective.com
- The Unintelligibility is Ours: Notes on Chain of Thought1a3orn.com
- Claude can make mistakes. Please double-check responses. | Julian Michaeljulianmichael.org
- Fake thinking and real thinking - Joe Carlsmithjoecarlsmith.com
- What Makes Art Great? - Nabeel S. Qureshinabeelqu.substack.com
- Short Story on AI: Forward Passkarpathy.github.io