✳flâneur — a map of the web's best reading
Quantifying Truesight With SAEs · Gwern.net
gwern.net · 1,991 words · saved by 2 readers
Proposal to use SAEs to crack open the dark matter of LLM inference about text.
--- title: Quantifying Truesight With SAEs description: "Proposal to use SAEs to crack open the dark matter of LLM inference about text." created: 2025-05-25 modified: 2025-05-26 status: finished importance: 7 confidence: possible css-extension: dropcaps-kanzlei ... > LLMs know more than they say, like who wrote some text. > How can we find out these known unknowns of how much they *really* know, when their internals are a mysterious scrambled blackbox? > > But we can possibly use sparse autoencoders to turn their internal thoughts into a long list of simpler properties summarizing what the LL
Explore this link on the map →saved by
related reading
- Writing for LLMs So They Listen · Gwern.netgwern.net
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- Natural Language Autoencoders \ Anthropicanthropic.com
- Neuronpedianeuronpedia.org
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com
- pdfopenreview.net
- The case for more ambitious language model evals — LessWronglesswrong.com