Oversight Assistants: Turning Compute into Understanding
Currently, we primarily oversee AI with human supervision and human-run experiments, possibly augmented by off-the-shelf AI assistants like ChatGPT or Claude. At training time, we run RLHF, where humans (and/or chat assistants) label behaviors with whether they are good or not. Afterwards, human researchers do additional testing to surface and evaluate unwanted behaviors, possibly assisted by a scaffolded chat agent. The problem with primarily human-driven oversight is that it is not scalable: as AI systems keep getting smarter, errors become harder to detect: For all these reasons, we need oversight mechanisms that scale beyond human overseers and that can grapple with the increasing sophistication of AI agents. Augmenting humans with current-generation chatbots does not resolve these issues: such off-the-shelf oversight won’t be superhuman until general AI systems are superhuman, which is too late. Moreover, capabilities are spiky—models can excel at some tasks while doing poorly at
Currently, we primarily oversee AI with human supervision and human-run experiments, possibly augmented by off-the-shelf AI assistants like ChatGPT or Claude. At training time, we run RLHF , where humans (and/or chat assistants ) label behaviors with whether they are good or not. Afterwards, human researchers do additional testing to surface and evaluate unwanted behaviors, possibly assisted by a scaffolded chat agent . The problem with primarily human-driven oversight is that it is not scalable: as AI systems keep getting smarter, errors become harder to detect: The behaviors we care about be
Explore this link on the map →saved by
related reading
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- AI 2027ai-2027.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org
- Foundation Models for Oversight | Transluce AItransluce.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Prediction, Explanation, or Over-interpretation?elena-baixy.github.io
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- The Hitchhiker's Guide to Actionable Interpretabilityactionable-interpretability-guide.github.io