To be legible, evidence of misalignment probably has to be behavioral
blog.redwoodresearch.org · 1,246 words · saved by 1 readers
Evidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing.
To be legible, evidence of misalignment probably has to be behavioral Evidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing. Ryan Greenblatt Apr 15, 2025 9 3 1 Share One key hope for mitigating risk from misalignment is inspecting the AI's behavior, noticing that it did something egregiously bad , converting this into legible evidence the AI is seriously misaligned, and then this triggering some strong and useful response (like spending relatively more resources on safety or undeploying this misaligned AI). You might hope that (fancy) internals-based t
saved by
related reading
- The Case for Model Forensics — LessWronglesswrong.com
- How independent researchers could investigate AI propensities after misalignment incidents - METRmetr.org
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- Not a Paper: "Frontier Lab CEOs are Capable of In-Context Scheming" — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- [2606.26071] Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignmentarxiv.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Incriminating misaligned AI models via distillation — LessWronglesswrong.com
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org