Driving fast in the counterfactual loop | Ordinary Ideas
The car is perfectly capable of driving safely. But what happens in the 1% of cases where the car decides to pause and ask the human to review its proposed behavior, or to suggest an action? Without some further precautions, the car is liable to immediately crash, and so the human won’t be able to provide any useful oversight at all. And that means that in the 99% of cases where the robot doesn’t ask the human for feedback, it won’t do anything useful. Foreseeing this outcome, the human may install a backup system to drive the car while the first system is suspended. Unfortunately this doesn’t fix the problem. If the first system pauses, then the backup could spring into action. But if it also pauses, then the car will crash. And so the second system won’t do anything useful if the first system pauses. And so the first system won’t do anything useful. As far as I can tell, no collection of counterfactually supervised systems can drive a car that contains the overseer. Of course that’s
An allegory Consider a human controlling a very fast car on a busy street using counterfactual oversight . The car is perfectly capable of driving safely. But what happens in the 1% of cases where the car decides to pause and ask the human to review its proposed behavior, or to suggest an action? Without some further precautions, the car is liable to immediately crash, and so the human won’t be able to provide any useful oversight at all. And that means that in the 99% of cases where the robot doesn’t ask the human for feedback, it won’t do anything useful. Foreseeing this outcome, the human m
Explore this link on the map →related reading
- What failure looks like — LessWronglesswrong.com
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org
- Towards self-driving codebases · Cursorcursor.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- How might we safely pass the buck to AI? — LessWronglesswrong.com
- Notes on handling non-concentrated failures with AI control: high level methods and different regimes — LessWronglesswrong.com
- Fields that I reference when thinking about AI takeover prevention — LessWronglesswrong.com
- AI Safety Seems Hard to Measurecold-takes.com