flâneur — a map of the web's best reading

Jankily controlling superintelligence - by Ryan Greenblatt

redwoodresearch.substack.com · 2,471 words · saved by 1 readers

When discussing AI control, we often talk about levels of AI capabilities where we think control can probably greatly lower risks and where we can probably estimate risks. However, I think it's plausible that an important application of control is modestly improving our odds of surviving significantly superhuman systems which are misaligned. This won't involve any real certainty, it would only moderately reduce risk (if it reduces risk at all), and it may only help with avoiding immediate loss of control (rather than also helping us extract useful work). In a sane world, we wouldn't be relying on control for systems which have significantly superhuman general purpose capabilities. Nevertheless, I think applying control here could be worthwhile. I'll discuss why I think this might work and how useful I think applying control in this way would be. As capabilities advance, the level of risk increases if we hold control measures fixed. There are a few different mechanisms that cause this,

Jankily controlling superintelligence How much time can control buy us during the intelligence explosion? Ryan Greenblatt Jun 27, 2025 8 Share When discussing AI control , we often talk about levels of AI capabilities where we think control can probably greatly lower risks and where we can probably estimate risks. However, I think it's plausible that an important application of control is modestly improving our odds of surviving significantly superhuman systems which are misaligned. This won't involve any real certainty, it would only moderately reduce risk (if it reduces risk at all), and it

Explore this link on the map →

related reading