Alignment will happen by default. What’s next? — LessWrong
I’m not 100% convinced of this, but I’m fairly convinced, more and more so over time. I’m hoping to start a vigorous but civilized debate. I invite y…
x Alignment will happen by default. What’s next? — LessWrong AI Frontpage 2025 Top Fifty: 7 % 98 Alignment will happen by default. What’s next? by Adrià Garriga-alonso 25th Nov 2025 AI Alignment Forum Linkpost for open.substack.com 7 min read 133 98 Ω 25 I’m not 100% convinced of this, but I’m fairly convinced, more and more so over time. I’m hoping to start a vigorous but civilized debate. I invite you to attack my weak points and/or present counter-evidence. My thesis is that intent-alignment is basically happening, based on evidence from the alignment research in the LLM era. Introduction T
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Alignment faking in large language modelsarxiv.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- the void — LessWronglesswrong.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org