Oversight Assistants: Turning Compute into Understanding
Currently, we primarily oversee AI with human supervision and human-run experiments, possibly augmented by off-the-shelf AI assistants like ChatGPT or Claude. At training time, we run RLHF, where humans (and/or chat assistants) label behaviors with whether they are good or not. Afterwards, human researchers do additional testing to surface
Currently, we primarily oversee AI with human supervision and human-run experiments, possibly augmented by off-the-shelf AI assistants like ChatGPT or Claude. At training time, we run RLHF , where humans (and/or chat assistants ) label behaviors with whether they are good or not. Afterwards, human researchers do additional testing to surface and evaluate unwanted behaviors, possibly assisted by a scaffolded chat agent . The problem with primarily human-driven oversight is that it is not scalable: as AI systems keep getting smarter, errors become harder to detect: The behaviors we care about be
related reading
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Measuring progress on scalable oversightanthropic.com
- Foundation Models for Oversight | Transluce AItransluce.org
- AI 2027ai-2027.com
- Scaling Activation Oracles to Trillion-Parameter Modelstransluce.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Measuring Progress on Scalable Oversight for Large Language Models \ Anthropicanthropic.com
- Can we scale human feedback for complex AI tasks? An intro to scalable oversight.aisafetyfundamentals.com
- Reflections On The Feasibility Of Scalable-Oversight — LessWronglesswrong.com
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org