A Summary of Recent Work (July 2026)
gdmalignment.substack.com · 2,613 words · saved by 2 readers
By Rohin Shah and Seb Farquhar.
It’s been nearly two years since our last major update in August 2024 and we wanted to share another recap of our recent work. Things have changed a lot since then. We are now fully in the midgame, and focus more on landing things in production. We are the AGI Safety and Alignment Team (ASAT), the main group at Google DeepMind working directly on technical approaches to existential risk from AI systems. Last year we published An Approach to Technical AGI Safety and Security, which remains the best place to read our overarching vision. Norms around chain of thought. Our impression is that…
saved by
related reading
- SOTA alignment assessments don’t strongly update us against misalignmentblog.redwoodresearch.org
- Claude Opus 4.5: Model Card, Alignment and Safetythezvi.substack.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Off Target | CNAScnas.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Rohin Shah on what it's really like to run AGI safety at Google DeepMind (and where I disagree with 'doomers') | 80,000 Hours80000hours.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- The Artificiality of Alignmentjoinreboot.org
- AI Safety | Arkosevictoriabrook.github.io