Evaluating Sparse Autoencoders with Board Games | Adam Karvonen
This blog post discusses a collaborative research paper on sparse autoencoders (SAEs), specifically focusing on SAE evaluations and a new training method we call p-annealing. As the first author, I primarily contributed to the evaluation portion of our work. The views expressed here are my own and do not necessarily reflect the perspectives of my co-authors. You can access our full paper here.
This blog post discusses a collaborative research paper on sparse autoencoders (SAEs), specifically focusing on SAE evaluations and a new training method we call p-annealing . As the first author, I primarily contributed to the evaluation portion of our work. The views expressed here are my own and do not necessarily reflect the perspectives of my co-authors. You can access our full paper here . Key Results In our research on evaluating Sparse Autoencoders (SAEs) using board games, we had several key findings: We developed two new metrics for evaluating Sparse Autoencoders (SAEs) in the contex
Explore this link on the map →saved by
related reading
- Research Report: Sparse Autoencoders find only 9/180 board state features in OthelloGPT — LessWronglesswrong.com
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- Interpretability with Sparse Autoencoders (Colab exercises) — LessWronglesswrong.com
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- Matryoshka Sparse Autoencoders — LessWronglesswrong.com
- pdfopenreview.net
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Do sparse autoencoders find "true features"? — LessWronglesswrong.com
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- Chapter 1: Transformer Interpretability - ARENAlearn.arena.education