flâneur — a map of the web's best reading

Catching AIs red-handed — LessWrong

lesswrong.com · 11,717 words · saved by 1 readers

We’ve been spending a lot of time recently thinking about how to mitigate risks posed by scheming (a.k.a. deceptively aligned) models intentionally s…

x Catching AIs red-handed — LessWrong Best of LessWrong 2024 Deceptive Alignment AI Control Redwood Research AI Frontpage 116 Catching AIs red-handed by ryan_greenblatt , Buck 5th Jan 2024 AI Alignment Forum 21 min read 28 116 Ω 62 We’ve been spending a lot of time recently thinking about how to mitigate risks posed by scheming (a.k.a. deceptively aligned) models intentionally subverting your safety techniques in order to cause an unacceptable outcome. We empirically investigated techniques that are robust to intentional subversion in our recent paper . In this post, we’ll discuss a crucial dy

Explore this link on the map →

related reading