flâneur — a map of the web's best reading

Mechanistic Transparency for Machine Learning

danielfilan.com · 1,148 words · saved by 1 readers

[EDIT (added Jan 2023): it’s come to my attention that this post was likely influenced by conversations I had with Chris Olah related to the distinction between standard interpretability and the type called “mechanistic”, as well as early experiments he had which became the ‘circuits’ sequence of papers - my sincere apologies for not making this clearer earlier.]

Mechanistic Transparency for Machine Learning Mechanistic Transparency for Machine Learning Jul 10, 2018 Cross-posted to the AI Alignment Forum . [EDIT (added Jan 2023): it’s come to my attention that this post was likely influenced by conversations I had with Chris Olah related to the distinction between standard interpretability and the type called “mechanistic”, as well as early experiments he had which became the ‘circuits’ sequence of papers - my sincere apologies for not making this clearer earlier.] Lately I’ve been trying to come up with a thread of AI alignment research that (a) I can

Explore this link on the map →

related reading