flâneur — a map of the web's best reading

Catching AIs red-handed

blog.redwoodresearch.org · 6,967 words · saved by 1 readers

If your AIs are trying to escape, it's crucial to think about whether you can catch them before they succeed, because catching them red-handed gives you lots of options you didn't have before.

Catching AIs red-handed If your AIs are trying to escape, it's crucial to think about whether you can catch them before they succeed, because catching them red-handed gives you lots of options you didn't have before. Buck Shlegeris and Ryan Greenblatt May 07, 2024 5 1 Share [Originally posted on LessWrong .] We’ve been spending a lot of time recently thinking about how to mitigate risks posed by scheming (a.k.a. deceptively aligned) models intentionally subverting your safety techniques in order to cause an unacceptable outcome. We empirically investigated techniques that are robust to intenti

Explore this link on the map →

saved by

related reading