A List of 45+ Mech Interp Project Ideas from Apollo Research’s Interpretability Team — AI Alignment Forum
We hope some people find this list helpful! We would love to see people working on these! If any sound interesting to you and you'd like to chat about it, don't hesitate to reach out. Papers from our first project here and here and from our second project here. Therefore, many project ideas in that list aren’t an up-to-date reflection of what some researchers consider the frontiers of mech interp. Can confirm, that list is SO out of date and does not represent the current frontiers. Zero offence taken. Thanks for publishing this list! recently[1]. empty footnote Thanks! Fixed now
x A List of 45+ Mech Interp Project Ideas from Apollo Research’s Interpretability Team — AI Alignment Forum Interpretability (ML & AI) Apollo Research (org) Sparse Autoencoders (SAEs) AI Frontpage 52 A List of 45+ Mech Interp Project Ideas from Apollo Research’s Interpretability Team by Lee Sharkey , Lucius Bushnaq , Dan Braun , StefanHex , Nicholas Goldowsky-Dill 18th Jul 2024 22 min read 18 52 Why we made this list: The interpretability team at Apollo Research wrapped up a few projects recently [1] . In order to decide what we’d work on next, we generated a lot of different potential project
Explore this link on the map →related reading
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- An Intuitive Explanation of Sparse Autoencoders for LLM Interpretability | Adam Karvonenadamkarvonen.github.io
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A gentle introduction to sparse autoencoders — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- pdfopenreview.net
- Sparse Autoencoders Find Highly Interpretable Features in Language Modelsarxiv.org
- Sparsify: A mechanistic interpretability research agenda — AI Alignment Forumalignmentforum.org
- A gentle introduction to sparse autoencodersnickjiang.substack.com
- An Intuitive Explanation of Sparse Autoencoders for Mechanistic Interpretability of LLMs — LessWronglesswrong.com
- Identifying Functionally Important Features with End-to-End Sparse Dictionary Learning — LessWronglesswrong.com