flâneur — a map of the web's best reading

Fact Finding: Simplifying the Circuit (Post 2) — LessWrong

lesswrong.com · 6,172 words · saved by 1 readers

This is the second post in the Google DeepMind mechanistic interpretability team’s investigation into how language models recall facts. This post foc…

x Fact Finding: Simplifying the Circuit (Post 2) — LessWrong Interpretability (ML & AI) Frontpage 27 Fact Finding: Simplifying the Circuit (Post 2) by Senthooran Rajamanoharan , Neel Nanda , János Kramár , Rohin Shah 23rd Dec 2023 AI Alignment Forum 17 min read 3 27 Ω 16 This is the second post in the Google DeepMind mechanistic interpretability team’s investigation into how language models recall facts . This post focuses on distilling down the fact recall circuit and models a more standard mechanistic interpretability investigation. This post gets in the weeds, we recommend starting with pos

Explore this link on the map →

related reading