flâneur — a map of the web's best reading

EIS XIV: Is mechanistic interpretability about to be practically useful? — AI Alignment Forum

alignmentforum.org · 2,864 words · saved by 1 readers

Lately, I have been thinking of interpretability research as falling into five different tiers of rigor. This is when researchers claim they have succeeded in interpreting a model by definition or based on analyzing results and asserting hypotheses about them. This is a key part of the scientific method. But by itself, it is not good science. Previously in this sequence, I have argued that this standard is fairly pervasive. This is when researchers develop an interpretation, use it to make some (usually simple) prediction, and then show that this prediction validates. This is at least doing science, but it doesn't necessarily demonstrate any usefulness or value. This is when researchers accomplish a useful type of task with an interpretability technique but do so in a way that is toy, cherry-picked, or under a streetlight. This is when researchers show that an interpretability tool can be used uniquely or competitively to accomplish a useful task. For this level of rigor, it needs to b

x EIS XIV: Is mechanistic interpretability about to be practically useful? — AI Alignment Forum The Engineer’s Interpretability Sequence Interpretability (ML & AI) AI Frontpage 31 EIS XIV: Is mechanistic interpretability about to be practically useful? by scasper 11th Oct 2024 8 min read 4 31 Part 14 of 12 in the Engineer’s Interpretability Sequence . Is this market really only at 63%? I think you should take the over. Only 63%? I think you should take the over. Five tiers of rigor for safety-oriented interpretability work Lately, I have been thinking of interpretability research as falling in

Explore this link on the map →

related reading