flâneur — a map of the web's best reading

Matryoshka Sparse Autoencoders — LessWrong

lesswrong.com · 6,247 words · saved by 2 readers

View trees here Search through latents with a token-regex language View individual latents here See code here (github.com/noanabeshima/matryoshka-sae…

x Matryoshka Sparse Autoencoders — LessWrong Interpretability (ML & AI) Sparse Autoencoders (SAEs) AI Frontpage 100 Matryoshka Sparse Autoencoders by Noa Nabeshima 14th Dec 2024 AI Alignment Forum 13 min read 15 100 Ω 51 View trees here Search through latents with a token-regex language View individual latents here See code here (github.com/noanabeshima/matryoshka-saes) Alternate version of this document with appropriate-height interactives. Abstract Sparse autoencoders (SAEs) [1] [2] break down neural network internals into components called latents. Smaller SAE latents seem to correspond to

Explore this link on the map →

saved by

related reading