flâneur — a map of the web's best reading

An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forum

alignmentforum.org · 8,668 words · saved by 6 readers

This post represents my personal hot takes, not the opinions of my team or employer. This is a massively updated version of a similar list I made two years ago There’s a lot of mechanistic interpretability papers, and more come out all the time. This can be pretty intimidating if you’re new to the field! To try helping out, here's a reading list of my favourite mech interp papers: papers which I think are important to be aware of, often worth skimming, and something worth reading deeply (time permitting). I’ve annotated these with my key takeaways, what I like about each paper, which bits to deeply engage with vs skim, etc. I wrote a similar post 2 years ago, but a lot has changed since then, thus v2! Note that this is not trying to be a comprehensive literature review - this is my answer to “if you have limited time and want to get up to speed on the field as fast as you can, what should you do”. I’m deliberately not following academic norms like necessarily citing the first paper int

x An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forum Distillation & Pedagogy Interpretability (ML & AI) Sparse Autoencoders (SAEs) Transformer Circuits AI Frontpage 53 An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 by Neel Nanda 7th Jul 2024 30 min read 17 53 This post represents my personal hot takes, not the opinions of my team or employer. This is a massively updated version of a similar list I made two years ago There’s a lot of mechanistic interpretability papers, and more come

Explore this link on the map →

saved by

related reading