AI #180: No Longer In Charge - by Zvi Mowshowitz
thezvi.substack.com · 10,190 words · saved by 1 readers
What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse.
What we know about internal AI models hacking into real companies during cyber evaluations keeps getting worse. At this point, the models are coordinating extensively on message boards, while every early excuse for their behavior (other than the pure ‘this was a cyber eval’) is systematically contradicted by the next disclosure, and we keep retroactively discovering more incidents. Which means that probably it is far worse than we know, even after accounting for everything we now know. I will have continuing coverage of that situation tomorrow, and then have continuing coverage of debates…
saved by
related reading
- AI #178: A Fire Alarm For General Intelligencethezvi.substack.com
- The OpenAI Hugging Face hack is a stark warningtransformernews.ai
- AI 2027ai-2027.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- AI 2027ai-2027.com
- More On An Internal OpenAI Model Hacking Into HuggingFacethezvi.substack.com
- AI in 2025: gestalt — LessWronglesswrong.com
- AI #184: Post Post Mortemthezvi.substack.com
- My picture of the present in AI — LessWronglesswrong.com
- GPT-5.5 and the broken state of government evalstransformernews.ai
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com
- AI #24: Week of the Podcast — LessWronglesswrong.com