constitutional.pdf
anthropic.com · 295 words · saved by 2 readers
N/A
Abstract Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on most skills relevant to the task at hand. Empirical work on this problem is not straightforward, since we do not yet have systems that broadly exceed our abilities. This paper discusses one of the major ways we think about this problem, with a focus on ways it can be studied empirically. We first present an experimental design centered on tasks for which human specialists succeed but unaided humans and…
saved by
related reading
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Measuring Progress on Scalable Oversight for Large Language Models \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Foundation Models for Oversight | Transluce AItransluce.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Scaling Activation Oracles to Trillion-Parameter Modelstransluce.org