Putting up Bumpers — LessWrong
tl;dr: Even if we can't solve alignment, we can solve the problem of catching and fixing misalignment. If a child is bowling for the first time, and they just aim at the pins and throw, they’re almost certain to miss. Their ball will fall into one of the gutters. But if there were beginners’ bumpers in place blocking much of the length of those gutters, their throw would be almost certain to hit at least a few pins. This essay describes an alignment strategy for early AGI systems I call ‘putting up bumpers’, in which we treat it as a top priority to implement and test safeguards that allow us to course-correct if we turn out to have built or deployed a misaligned model, in the same way that bowling bumpers allow a poorly aimed ball to reach its target. To do this, we'd aim to build up many largely-independent lines of defense that allow us to catch and respond to early signs of misalignment. This would involve significantly improving and scaling up our use of tools like mechanistic int
x Putting up Bumpers — LessWrong AI Auditing AI Control Anthropic (org) AI Frontpage 58 Putting up Bumpers by Sam Bowman 23rd Apr 2025 AI Alignment Forum 2 min read 14 58 Ω 31 tl;dr: Even if we can't solve alignment, we can solve the problem of catching and fixing misalignment. If a child is bowling for the first time, and they just aim at the pins and throw, they’re almost certain to miss. Their ball will fall into one of the gutters. But if there were beginners’ bumpers in place blocking much of the length of those gutters, their throw would be almost certain to hit at least a few pins. This
Explore this link on the map →related reading
- Putting up Bumpersalignment.anthropic.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- A minimal viable product for alignment - by Jan Leikealigned.substack.com
- Can we safely automate alignment research? - Joe Carlsmithjoecarlsmith.com
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Why I’m optimistic about our alignment approachaligned.substack.com