✳flâneur — a map of the web's best reading
The Claude Bliss Attractor - by Scott Alexander
astralcodexten.com · 1,229 words · saved by 1 readers
...
The Claude Bliss Attractor ... Scott Alexander Jun 13, 2025 532 302 35 Share This is a reported phenomenon where if two copies of Claude talk to each other, they end up spiraling into rapturous discussion of spiritual bliss, Buddhism, and the nature of consciousness. From the system card : Anthropic swears they didn’t do this on purpose; when they ask Claude why this keeps happening, Claude can’t explain. Needless to say, this has made lots of people freak out / speculate wildly. I think there are already a few good partial explanations of this (especially Nostalgebraist here ), but they deser
Explore this link on the map →related reading
- Child’s Play, by Sam Krissharpers.org
- models have some pretty funny attractor states — LessWronglesswrong.com
- When AI builds itself \ Anthropicanthropic.com
- Claude’s Character \ Anthropicanthropic.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- What Is Claude? Anthropic Doesn’t Know, Either | The New Yorkernewyorker.com
- Claude’s Constitution \ Anthropicanthropic.com
- A global workspace in language models \ Anthropicanthropic.com
- Claude’s Character \ Anthropicanthropic.com
- Claude's Character | Hacker Newsnews.ycombinator.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com