AI #178: A Fire Alarm For General Intelligence
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym. It is much more important that you read those two posts, and the one on Kimi K3, than to read this one that rounds up the other news of the week. OpenAI wants to present this as largely an infrastructure and safeguards problem, that it needs to build more secure sandboxes and have better…
saved by
related reading
- The OpenAI Hugging Face hack is a stark warningtransformernews.ai
- AI #180: No Longer In Chargethezvi.substack.com
- AI 2027ai-2027.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- More On An Internal OpenAI Model Hacking Into HuggingFacethezvi.substack.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?blog.redwoodresearch.org
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?blog.redwoodresearch.org
- AI Lab Watchailabwatch.org