flâneur — a map of the web's best reading

On the functional self of LLMs — LessWrong

lesswrong.com · 10,561 words · saved by 1 readers

Anthropic's 'Scaling Monosemanticity' paper got lots of well-deserved attention for its work taking sparse autoencoders to a new level. But I was absolutely transfixed by a short section near the end, 'Features Relating to the Model’s Representation of Self', which explores what SAE features activate when the model is asked about itself[1]: Some of those features are reasonable representations of the assistant persona — but some of them very much aren't. 'Spiritual beings like ghosts, souls, or angels'? 'Artificial intelligence becoming self-aware'? What's going on here? The authors 'urge caution in interpreting these results…How these features are used by the model remains unclear.' That seems very reasonable — but how could anyone fail to be intrigued? Seeing those results a year ago started me down the road of asking what we can say about what LLMs believe about themselves, how that connects to their actual values and behavior, and how shaping their self-model could help us build be

x On the functional self of LLMs — LessWrong Interpretability (ML & AI) Language Models (LLMs) RLHF Situational Awareness AI Frontpage 2025 Top Fifty: 11 % 124 On the functional self of LLMs by eggsyntax 7th Jul 2025 AI Alignment Forum 10 min read 38 124 Ω 45 Summary Introduces a research agenda I believe is important and neglected: investigating whether frontier LLMs acquire something functionally similar to a self, a deeply internalized character with persistent values, outlooks, preferences, and perhaps goals; exploring how that functional self emerges; understanding how it causally interac

Explore this link on the map →

related reading