Oversight Assistants: Turning Compute into Understanding
Currently, we primarily oversee AI with human supervision and human-run experiments, possibly augmented by off-the-shelf AI assistants like ChatGPT or Claude. At training time, we run RLHF, where humans (and/or chat assistants) label behaviors with whether they are good or not. Afterwards, human researchers do additional testing to surface
Currently, we primarily oversee AI with human supervision and human-run experiments, possibly augmented by off-the-shelf AI assistants like ChatGPT or Claude. At training time, we run RLHF , where humans (and/or chat assistants ) label behaviors with whether they are good or not. Afterwards, human researchers do additional testing to surface and evaluate unwanted behaviors, possibly assisted by a scaffolded chat agent . The problem with primarily human-driven oversight is that it is not scalable: as AI systems keep getting smarter, errors become harder to detect: The behaviors we care about be
Explore this link on the map →related reading
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- AI 2027ai-2027.com
- Foundation Models for Oversight | Transluce AItransluce.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Measuring Progress on Scalable Oversight for Large Language Models \ Anthropicanthropic.com
- Reflections On The Feasibility Of Scalable-Oversight — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org
- AI 2027ai-2027.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com