flâneur

Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?

blog.redwoodresearch.org · 1,638 words · saved by 1 readers

Yes, but less than had they been schemers.

OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted1. Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions. We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis. The myopic, unambitious misalignment…

saved by

related reading