Alignment is not solved but it increasingly looks solvable
aligned.substack.com · 1,490 words · saved by 17 readers
But it increasingly looks solvable
I’ve been optimistic about alignment for a while now, but when I first wrote about this in 2022, there was a lot more uncertainty about how the technology would develop. Since then a lot has happened: pretraining continued improving and RL became a much bigger deal. A priori it wasn’t obvious that we can robustly align LLMs through the RL scale-up, because some alignment threat models are about models that become agentic, learn to pursue unaligned instrumental goals, and become deceptive in the process. In fact, the early highly RL’ed models like o1, o3, and Claude 3.7 exhibit a number of…
saved by
- Yixiong Hao
- Uzay Girit
- Aaron Pham
- Sudarsh K
- Vincent Cheng
- Yudhister Joel Kumar
- Eric Huang
- Jason Hausenloy
- Ishan Mukherjee
- Christine Ye
- Will Anderson
- Kunvar Thaman
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Teaching Claude Whyalignment.anthropic.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- SOTA alignment assessments don’t strongly update us against misalignmentblog.redwoodresearch.org
- Teaching Claude why \ Anthropicanthropic.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Why I’m optimistic about our alignment approachaligned.substack.com