flâneur — a map of the web's best reading

On Developing a Mathematical Theory of Interpretability — LessWrong

lesswrong.com · 7,104 words · saved by 1 readers

If the trajectory of the deep learning paradigm continues, it seems plausible to me that in order for applications of low-level interpretability to AI not-kill-everyone-ism to be truly reliable, we will need a much better-developed and more general theoretical and mathematical framework for deep learning than currently exists. And this sort of work seems difficult. Doing mathematics carefully - in particular finding correct, rigorous statements and then finding correct proofs of those statements - is slow. So slow that the rate of change of cutting-edge engineering practices significantly worsens the difficulties involved in building theory at the right level of generality. And, in my opinion, much slower than the rate at which we can generate informal observations that might possibly be worthy of further mathematical investigation. Thus it can feel like the role that serious mathematics has to play in interpretability is primarily reactive, i.e. consists mostly of activities like 'add

x On Developing a Mathematical Theory of Interpretability — LessWrong Interpretability (ML & AI) Logic & Mathematics AI Frontpage 64 On Developing a Mathematical Theory of Interpretability by carboniferous_umbraculum 9th Feb 2023 AI Alignment Forum 8 min read 8 64 Ω 31 If the trajectory of the deep learning paradigm continues, it seems plausible to me that in order for applications of low-level interpretability to AI not-kill-everyone-ism to be truly reliable, we will need a much better-developed and more general theoretical and mathematical framework for deep learning than currently exists. A

Explore this link on the map →

saved by

related reading