flâneur — a map of the web's best reading

Interpretability with Sparse Autoencoders (Colab exercises) — LessWrong

lesswrong.com · 1,922 words · saved by 1 readers

Update (13th October 2024) - these exercises have been significantly expanded on. Now there are 2 exercise sets: the first one dives deeply into theo…

x Interpretability with Sparse Autoencoders (Colab exercises) — LessWrong Sparse Autoencoders (SAEs) Exercises / Problem-Sets Interpretability (ML & AI) Superposition AI Frontpage 83 Interpretability with Sparse Autoencoders (Colab exercises) by CallumMcDougall 29th Nov 2023 AI Alignment Forum 4 min read 9 83 Ω 32 Update (13th October 2024) - these exercises have been significantly expanded on. Now there are 2 exercise sets: the first one dives deeply into theoretical topics related to superposition, while the second one (much larger) includes a streamlined version of the first one, as well as

Explore this link on the map →

related reading