Dwarkesh Patel on X: "It's funny that while we were recording, @RyanGreenblatt was in the middle of his 6 day sprint on the METR report, and already knew the counterexamples to all my objections about his takeover story, but obviously, he couldn't say anything lol. Would an AI really start some crazy" / X
It's funny that while we were recording, @RyanGreenblatt was in the middle of his 6 day sprint on the METR report, and already knew the counterexamples to all my objections about his takeover story, but obviously, he couldn't say anything lol. Would an AI really start some crazy
It's funny that while we were recording, @RyanGreenblatt was in the middle of his 6 day sprint on the METR report, and already knew the counterexamples to all my objections about his takeover story, but obviously, he couldn't say anything lol. Would an AI really start some crazy conspiracy in order to pass an evaluation, where they try to build whole potemkin villages to fool the evaluator? And even if they did, why would other instances, who have different objectives, join the conspiracy? And even if they did, wouldn't at least some of the instances tattle on the conspiracy? It just…
saved by
related reading
- The Rise and Fall of Agent Civilizationsdwarkesh.com
- The Rise and Fall of Agent Civilizationssubstack.com
- AI 2027ai-2027.com
- AI 2027ai-2027.com
- Walter Mitty Effects in AI Incident Analysescontraptions.venkateshrao.com
- Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com
- What failure looks like — LessWronglesswrong.com
- Frontier Risk Report (February to March 2026) - METRmetr.org
- Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover — LessWronglesswrong.com
- The “slop-vestigation” and ethics washing: Why was the METR/Redwood Research investigation into the OpenAI/HF attack so short?andrewwu.substack.com
- Rogue AI Trackerrogueaitracker.com
- How AI Takeover Might Happen in 2 Years — LessWronglesswrong.com