flâneur — a map of the web's best reading

Measuring Progress on Scalable Oversight for Large Language Models \ Anthropic

anthropic.com · 215 words · saved by 1 readers

Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.

Alignment Research Measuring Progress on Scalable Oversight for Large Language Models Nov 4, 2022 Read Paper Abstract Developing safe and useful general-purpose AI systems will require us to make progress on scalable oversight: the problem of supervising systems that potentially outperform us on most skills relevant to the task at hand. Empirical work on this problem is not straightforward, since we do not yet have systems that broadly exceed our abilities. This paper discusses one of the major ways we think about this problem, with a focus on ways it can be studied empirically. We first prese

Explore this link on the map →

related reading