✳flâneur — a map of the web's best reading
Reflections on Neuralese — LessWrong
lesswrong.com · 1,965 words · saved by 1 readers
Interpretability is difficult, and making reasoning opaque makes it more difficult.
x Reflections on Neuralese — LessWrong Interpretability (ML & AI) AI Frontpage 48 Reflections on Neuralese by Alice Blair 12th Mar 2025 6 min read 3 48 Thanks to Brendan Halstead for feedback on an early draft of this piece. Any mistakes here are my own. [Epistemic status: I've looked at the relevant code enough to be moderately sure I understand what's going on. Predictions about the future, including about what facts will turn out to be relevant, are uncertain as always.] Introduction With the recent breakthroughs taking advantage of extensive Chain of Thought (CoT) reasoning in LLMs, there
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- Natural Language Autoencoders \ Anthropicanthropic.com
- 13 Arguments About a Transition to Neuralese AIs — LessWronglesswrong.com
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- How AI Is Learning to Think in Secret — LessWronglesswrong.com
- Language models can explain neurons in language modelsopenaipublic.blob.core.windows.net
- An idea for avoiding neuralese architectures — LessWronglesswrong.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWronglesswrong.com