SOTA alignment assessments don’t strongly update us against misalignment
Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1][2], it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”
Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherently misaligned[1][2], it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”[3]). While I agree with the report on the above bottom-line conclusions (substantially on priors)[4], I think there are gaps in its argument which weaken the current assessment and might invalidate future assessments. In particular,…
saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Claude Opus 4.5: Model Card, Alignment and Safetythezvi.substack.com
- A Summary of Recent Work (July 2026)gdmalignment.substack.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Claude Mythos knows when it's breaking the rules — and tries to hide itsubstack.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com