Papers and Projects - Josh Engels
joshengels.com · 405 words · saved by 1 readers
Publications on mechanistic interpretability and AI Safety.
Papers Paper | Blog | Twitter Janos Kramar, Joshua Engels, Zhengxuan Wang, Bilal Chughtai, Rohin Shah, Neel Nanda, and Arthur Conmy.Paper | Twitter Reilly Haskins, Bilal Chughtai, and Joshua Engels.Paper | Blog | Twitter Daria Ivanova, Riya Tyagi, Joshua Engels, and Neel Nanda.Paper | Blog Aria Wong, Joshua Engels, and Neel Nanda.Paper | Blog Joshua Engels*, David Baek*, Subhash Kantamneni*, and Max Tegmark.Paper | Code | Twitter Subhash Kantamneni*, Joshua Engels*, Senthooran Rajamanoharan, Max Tegmark, and Neel Nanda.Paper | Blog | Code | Twitter Mathew Chen*, Joshua Engels*, and…
saved by
related reading
- Transformer Circuits Threadtransformer-circuits.pub
- ARENA - AI Safety Curriculumlearn.arena.education
- Assessing skeptical views of interpretability research | Christopher Pottsweb.stanford.edu
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- An Extremely Opinionated Annotated List of My Favourite Mechanistic Interpretability Papers v2 — AI Alignment Forumalignmentforum.org
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org