My AGI safety research—2025 review, ’26 plans — AI Alignment Forum
“Our greatest fear should not be of failure, but of succeeding at something that doesn't really matter.” –attributed to DL Moody[1] The main threat model I’m working to address is the same as it’s been since I was hobby-blogging about AGI safety in 2019. Basically, I think that: I think that, when this learning algorithm is understood, it will be easy to get it to do powerful and impressive things, and to make money, as long as it’s weak enough that humans can keep it under control. But past that stage, we’ll be relying on the AGIs to have good motivations, and not be egregiously misaligned and scheming to take over the world and wipe out humanity. Alas, I claim that the latter kind of motivation is what we should expect to occur, in the absence of yet-to-be-invented techniques to avoid it. Inventing those yet-to-be-invented techniques constitutes the technical alignment problem for brain-like AGI. That’s the main thing I’ve been working on since I’ve been in the field. See my Intro to
x My AGI safety research—2025 review, ’26 plans — AI Alignment Forum Research Agendas AI Frontpage 2025 Top Fifty: 9 % 47 My AGI safety research—2025 review, ’26 plans by Steven Byrnes 11th Dec 2025 15 min read 4 47 Previous: 2024 , 2022 “Our greatest fear should not be of failure, but of succeeding at something that doesn't really matter.” – attributed to DL Moody [1] 1. Background & threat model The main threat model I’m working to address is the same as it’s been since I was hobby-blogging about AGI safety in 2019. Basically, I think that: The “secret sauce” of human intelligence is a big u
Explore this link on the map →related reading
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- LessWronglesswrong.com
- My AGI safety research—2024 review, ’25 plans — AI Alignment Forumalignmentforum.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI safetysjbyrnes.com
- AGI safety career advice — EA Forumforum.effectivealtruism.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI in 2025: gestalt — LessWronglesswrong.com
- [Intro to brain-like-AGI safety] 1. What's the problem & Why work on it now? — AI Alignment Forumalignmentforum.org
- Planning for AGI and beyond | OpenAIopenai.com
- (My understanding of) What Everyone in Technical Alignment is Doing and Why — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com