Alignment will happen by default. What’s next? — LessWrong
I’m not 100% convinced of this, but I’m fairly convinced, more and more so over time. I’m hoping to start a vigorous but civilized debate. I invite y…
x Alignment will happen by default. What’s next? — LessWrong AI Frontpage 2025 Top Fifty: 7 % 98 Alignment will happen by default. What’s next? by Adrià Garriga-alonso 25th Nov 2025 AI Alignment Forum Linkpost for open.substack.com 7 min read 133 98 Ω 25 I’m not 100% convinced of this, but I’m fairly convinced, more and more so over time. I’m hoping to start a vigorous but civilized debate. I invite you to attack my weak points and/or present counter-evidence. My thesis is that intent-alignment is basically happening, based on evidence from the alignment research in the LLM era. Introduction T
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Teaching Claude Whyalignment.anthropic.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Simulated Users & Sad LLMs1a3orn.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- How confessions can keep language models honest | OpenAIopenai.com