On Developing a Mathematical Theory of Interpretability — LessWrong
If the trajectory of the deep learning paradigm continues, it seems plausible to me that in order for applications of low-level interpretability to AI not-kill-everyone-ism to be truly reliable, we will need a much better-developed and more general theoretical and mathematical framework for deep learning than currently exists. And this sort of work seems difficult. Doing mathematics carefully - in particular finding correct, rigorous statements and then finding correct proofs of those statements - is slow. So slow that the rate of change of cutting-edge engineering practices significantly worsens the difficulties involved in building theory at the right level of generality. And, in my opinion, much slower than the rate at which we can generate informal observations that might possibly be worthy of further mathematical investigation. Thus it can feel like the role that serious mathematics has to play in interpretability is primarily reactive, i.e. consists mostly of activities like 'add
x On Developing a Mathematical Theory of Interpretability — LessWrong Interpretability (ML & AI) Logic & Mathematics AI Frontpage 64 On Developing a Mathematical Theory of Interpretability by carboniferous_umbraculum 9th Feb 2023 AI Alignment Forum 8 min read 8 64 Ω 31 If the trajectory of the deep learning paradigm continues, it seems plausible to me that in order for applications of low-level interpretability to AI not-kill-everyone-ism to be truly reliable, we will need a much better-developed and more general theoretical and mathematical framework for deep learning than currently exists. A
saved by
related reading
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- What is the purpose of interpretability?ericjmichaud.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- On Optimism for Interpretabilitygoodfire.ai
- An Ambitious Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Assessing skeptical views of interpretability research | Christopher Pottsweb.stanford.edu
- Statistical Physics for Ambitious Interpretability: A Workshop Retrospective — LessWronglesswrong.com
- The necessity of machine learning theory in mitigating AI riskmishabelkin.substack.com
- [2604.21691] There Will Be a Scientific Theory of Deep Learningarxiv.org
- A Pragmatic Vision for Interpretability — LessWronglesswrong.com
- Against Almost Every Theory of Impact of Interpretability — LessWronglesswrong.com
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org