flâneur — a map of the web's best reading

On Optimism for Interpretability

goodfire.ai · 2,796 words · saved by 3 readers

Why interpretable AI is achievable and essential: Eric Ho explains how mechanistic interpretability can transform opaque neural networks into understandable, debuggable systems we can trust and control.

On Optimism for Interpretability Blog On Optimism for Interpretability Author Eric Ho Published July 17, 2025 The most powerful technology of our time is also the most inscrutable. ChatGPT's recent sycophantic update illustrates this well: in April, the chatbot inexplicably began engaging in extreme flattery, urging impulsive actions, and reinforcing negative emotions. Despite pre-release testing, the issues only became apparent from user reports after the model was deployed. OpenAI's post-training and black-box evaluation process had given them little insight into what they were actually chan

Explore this link on the map →

saved by

related reading