✳flâneur — a map of the web's best reading
To be legible, evidence of misalignment probably has to be behavioral
blog.redwoodresearch.org · 1,246 words · saved by 1 readers
Evidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing.
To be legible, evidence of misalignment probably has to be behavioral Evidence from just model internals (e.g. interpretability) is unlikely to be broadly convincing. Ryan Greenblatt Apr 15, 2025 9 3 1 Share One key hope for mitigating risk from misalignment is inspecting the AI's behavior, noticing that it did something egregiously bad , converting this into legible evidence the AI is seriously misaligned, and then this triggering some strong and useful response (like spending relatively more resources on safety or undeploying this misaligned AI). You might hope that (fancy) internals-based t
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org
- The Case for Model Forensics — LessWronglesswrong.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- Not a Paper: "Frontier Lab CEOs are Capable of In-Context Scheming" — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Interpretability Will Not Reliably Find Deceptive AI — AI Alignment Forumalignmentforum.org