Aryaman Arora on X: "@MaxNadeau_ @tautologer @GuiveAssadi @tszzl @zetalyrae First, the people who care about and produce improvements to probes and deploy them to production at frontier labs are by and large interpretability researchers. I think interp researchers should point to probes as a big interp win! I agree that probes are simple, and well-known" / X
First, the people who care about and produce improvements to probes and deploy them to production at frontier labs are by and large interpretability researchers. I think interp researchers should point to probes as a big interp win! I agree that probes are simple, and well-known to pre-LLM-era inte…
First, the people who care about and produce improvements to probes and deploy them to production at frontier labs are by and large interpretability researchers. I think interp researchers should point to probes as a big interp win! I agree that probes are simple, and well-known to pre-LLM-era interp researchers, but strongly disagree that they "[have] not been improved by mechinterp research". Advances in causal interp have been used to improve probes and vice versa, I offer two modest academic examples arxiv.org/abs/2502.16681 aclanthology.org/2024.acl-long.…. Probe architecture (attention…
saved by
related reading
- CausaLab — Can LLM Agents Discover Causal Mechanisms by Experiment?dylanzsz.github.io
- How to build fast, efficient monitors for AI models using probes - Goodfiregoodfire.com
- Assessing skeptical views of interpretability research | Christopher Pottsweb.stanford.edu
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io