flâneur — a map of the web's best reading

Interpretability — LessWrong

lesswrong.com · 4,760 words · saved by 1 readers

Chris Olah wrote the following topic prompt for the Open Phil 2021 request for proposals on the alignment of AI systems. We (Asya Bergal and Nick Bec…

x Interpretability — LessWrong Open Philanthropy 2021 AI Alignment RFP Interpretability (ML & AI) AI Frontpage 61 Interpretability by abergal , Nick_Beckstead 29th Oct 2021 AI Alignment Forum 14 min read 13 61 Ω 33 Chris Olah wrote the following topic prompt for the Open Phil 2021 request for proposals on the alignment of AI systems. We (Asya Bergal and Nick Beckstead) are running the Open Phil RFP and are posting each section as a sequence on the Alignment Forum. Although Chris wrote this document, we didn’t want to commit him to being responsible for responding to comments on it by posting i

Explore this link on the map →

related reading