Anthropic announces interpretability advances. How much does this advance alignment? — LessWrong
Anthropic just published a pretty impressive set of results in interpretability. This raises for me, some questions and a concern: Interpretability h…
x Anthropic announces interpretability advances. How much does this advance alignment? — LessWrong Interpretability (ML & AI) AI Frontpage 49 Anthropic announces interpretability advances. How much does this advance alignment? by Seth Herd 21st May 2024 Linkpost for www.anthropic.com 4 min read 4 49 Anthropic just published a pretty impressive set of results in interpretability. This raises for me, some questions and a concern: Interpretability helps, but it isn't alignment, right? It seems to me as though the vast bulk of alignment funding is now going to interpretability. Who is thinking abo
Explore this link on the map →related reading
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Against Almost Every Theory of Impact of Interpretability — AI Alignment Forumalignmentforum.org
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Against Almost Every Theory of Impact of Interpretability — LessWronglesswrong.com