Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?
blog.redwoodresearch.org · 1,638 words · saved by 1 readers
Yes, but less than had they been schemers.
OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted1. Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions. We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis. The myopic, unambitious misalignment…
saved by
related reading
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?blog.redwoodresearch.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- AI #178: A Fire Alarm For General Intelligencethezvi.substack.com
- The Huggingface Incident - by Scott Alexanderastralcodexten.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- Training a Misaligned Reward Seekeralignment.anthropic.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com