Oversight Assistants: Turning Compute into Understanding
Currently, we primarily oversee AI with human supervision and human-run experiments, possibly augmented by off-the-shelf AI assistants like ChatGPT or Claude. At training time, we run RLHF, where humans (and/or chat assistants) label behaviors with whether they are good or not. Afterwards, human researchers do additional testing to surface and evaluate unwanted behaviors, possibly assisted by a scaffolded chat agent. The problem with primarily human-driven oversight is that it is not scalable: as AI systems keep getting smarter, errors become harder to detect: For all these reasons, we need oversight mechanisms that scale beyond human overseers and that can grapple with the increasing sophistication of AI agents. Augmenting humans with current-generation chatbots does not resolve these issues: such off-the-shelf oversight won’t be superhuman until general AI systems are superhuman, which is too late. Moreover, capabilities are spiky—models can excel at some tasks while doing poorly at
Currently, we primarily oversee AI with human supervision and human-run experiments, possibly augmented by off-the-shelf AI assistants like ChatGPT or Claude. At training time, we run RLHF , where humans (and/or chat assistants ) label behaviors with whether they are good or not. Afterwards, human researchers do additional testing to surface and evaluate unwanted behaviors, possibly assisted by a scaffolded chat agent . The problem with primarily human-driven oversight is that it is not scalable: as AI systems keep getting smarter, errors become harder to detect: The behaviors we care about be
saved by
related reading
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Measuring progress on scalable oversightanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Foundation Models for Oversight | Transluce AItransluce.org
- AI 2027ai-2027.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Scaling Activation Oracles to Trillion-Parameter Modelstransluce.org
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com