Measuring Progress on Scalable Oversight for Large Language Models \ Anthropic
anthropic.com · 215 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Alignment Research Measuring Progress on Scalable Oversight for Large Language Models Nov 4, 2022 Read Paper Abstract Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on most skills relevant to the task at hand. Empirical work on this problem is not straightforward, since we do not yet have systems that broadly exceed our abilities. This paper discusses one of the major ways we think about this problem, with a focus on ways it can be studied empirically. We first prese
related reading
- Measuring progress on scalable oversightanthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Anthropic's leading researchers acted as moderate accelerationists — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Foundation Models for Oversight | Transluce AItransluce.org