Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research (GDM Mech Interp Team Progress Update #2) — AI Alignment Forum
alignmentforum.org · 12,301 words · saved by 1 readers
Lewis Smith*, Sen Rajamanoharan*, Arthur Conmy, Callum McDougall, Janos Kramar, Tom Lieberum, Rohin Shah, Neel Nanda • * = equal contribution …
x Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research (GDM Mech Interp Team Progress Update #2) — AI Alignment Forum Interpretability (ML & AI) Sparse Autoencoders (SAEs) AI Curated 2025 Top Fifty: 39 % 58 Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research (GDM Mech Interp Team Progress Update #2) by lewis smith , Senthooran Rajamanoharan , Arthur Conmy , CallumMcDougall , Tom Lieberum , János Kramár , Rohin Shah , Neel Nanda 26th Mar 2025 Linkpost for deepmindsafetyresearch.medium.com 35 min read 15 58 Lewis Smith*, Sen Rajamanoharan*, Arth
related reading
- Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research (GDM Mech Interp Team Progress Update #2) — LessWronglesswrong.com
- Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research (GDM Mech Interp Team Progress Update #2) — AI Alignment Forumalignmentforum.org
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- pdfopenreview.net
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Modelsarxiv.org
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible.github.com
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com