flâneur — a map of the web's best reading

Reflections on Neuralese — LessWrong

lesswrong.com · 1,965 words · saved by 1 readers

Interpretability is difficult, and making reasoning opaque makes it more difficult.

x Reflections on Neuralese — LessWrong Interpretability (ML & AI) AI Frontpage 48 Reflections on Neuralese by Alice Blair 12th Mar 2025 6 min read 3 48 Thanks to Brendan Halstead for feedback on an early draft of this piece. Any mistakes here are my own. [Epistemic status: I've looked at the relevant code enough to be moderately sure I understand what's going on. Predictions about the future, including about what facts will turn out to be relevant, are uncertain as always.] Introduction With the recent breakthroughs taking advantage of extensive Chain of Thought (CoT) reasoning in LLMs, there

Explore this link on the map →

related reading