flâneur — a map of the web's best reading

Anthropic Fall 2023 Debate Progress Update — AI Alignment Forum

alignmentforum.org · 7,663 words · saved by 1 readers

This is a research update on some work that I’ve been doing on Scalable Oversight at Anthropic, based on the original AI safety via debate proposal a…

x Anthropic Fall 2023 Debate Progress Update — AI Alignment Forum Debate (AI safety technique) World Modeling Frontpage 32 Anthropic Fall 2023 Debate Progress Update by Ansh Radhakrishnan 28th Nov 2023 15 min read 9 32 This is a research update on some work that I’ve been doing on Scalable Oversight at Anthropic, based on the original AI safety via debate proposal and a more recent agenda developed at NYU and Anthropic. The core doc was written several months ago, so some of it is likely outdated, but it seemed worth sharing in its current form. I’d like to thank Tamera Lanham, Sam Bowman, Kam

Explore this link on the map →

related reading