flâneur — a map of the web's best reading

Against Almost Every Theory of Impact of Interpretability — LessWrong

lesswrong.com · 24,489 words · saved by 1 readers

Charbel-Raphaël argues that interpretability research has poor theories of impact. It's not good for predicting future AI systems, can't actually aud…

x Against Almost Every Theory of Impact of Interpretability — LessWrong Best of LessWrong 2023 Interpretability (ML & AI) AI Frontpage 336 Against Almost Every Theory of Impact of Interpretability by Charbel-Raphaël 17th Aug 2023 AI Alignment Forum 31 min read 93 336 Ω 96 Epistemic Status: I believe I am well-versed in this subject. I erred on the side of making claims that were too strong and allowing readers to disagree and start a discussion about precise points rather than trying to edge-case every statement. I also think that using memes is important because safety ideas are boring and an

Explore this link on the map →

related reading