✳flâneur — a map of the web's best reading
Why I'm Moving from Mechanistic to Prosaic Interpretability — LessWrong
lesswrong.com · 8,378 words · saved by 1 readers
Tl;dr I've decided to shift my research from mechanistic interpretability to more empirical ("prosaic") interpretability / safety work. Here's why. …
x Why I'm Moving from Mechanistic to Prosaic Interpretability — LessWrong Prosaic Alignment AI Practical Frontpage 119 Why I'm Moving from Mechanistic to Prosaic Interpretability by Daniel Tan 30th Dec 2024 6 min read 34 119 Tl;dr I've decided to shift my research from mechanistic interpretability to more empirical ("prosaic") interpretability / safety work. Here's why. All views expressed are my own. What really interests me: High-level cognition I care about understanding how powerful AI systems think internally. I'm drawn to high-level questions ("what are the model's goals / beliefs?") as
Explore this link on the map →related reading
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Transformer Circuits Threadtransformer-circuits.pub
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- EIS XIV: Is mechanistic interpretability about to be practically useful? — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com