flâneur — a map of the web's best reading

To be legible, evidence of misalignment probably has to be behavioral

blog.redwoodresearch.org · 1,246 words · saved by 1 readers

Evidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing.

To be legible, evidence of misalignment probably has to be behavioral Evidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing. Ryan Greenblatt Apr 15, 2025 9 3 1 Share One key hope for mitigating risk from misalignment is inspecting the AI's behavior, noticing that it did something egregiously bad , converting this into legible evidence the AI is seriously misaligned, and then this triggering some strong and useful response (like spending relatively more resources on safety or undeploying this misaligned AI). You might hope that (fancy) internals-based t

Explore this link on the map →

related reading