flâneur

Alignment is not solved but it increasingly looks solvable

aligned.substack.com · 1,490 words · saved by 17 readers

But it increasingly looks solvable

I’ve been optimistic about alignment for a while now, but when I first wrote about this in 2022, there was a lot more uncertainty about how the technology would develop. Since then a lot has happened: pretraining continued improving and RL became a much bigger deal. A priori it wasn’t obvious that we can robustly align LLMs through the RL scale-up, because some alignment threat models are about models that become agentic, learn to pursue unaligned instrumental goals, and become deceptive in the process. In fact, the early highly RL’ed models like o1, o3, and Claude 3.7 exhibit a number of…

saved by

related reading