Integrity incidents/issues/imperfections - AI Lab Watch
This page is not exhaustive, especially for labs besides OpenAI and Anthropic. Let us know of missing stuff. This page is from an x-risk perspective but includes things not directly relevant to x-risk. This page may include not just (central) integrity issues but also policy failures and statements that turned out to be very misleading. This page would include instances of high-integrity-ness but those are weird and hard to notice. Note that public communication is very good but leads to integrity incidents. More context: I was pulled aside for a chat with a lawyer that quickly turned adversarial. The questions were about my views on AI progress, on AGI, the appropriate level of security for AGI, whether the government should be involved in AGI, whether I and the superalignment team were loyal to the company, and what I was up to during the OpenAI board events. They then talked to a couple of my colleagues and came back and told me I was fired. They’d gone through all of my digital art
This page is not exhaustive, especially for companies besides OpenAI and Anthropic. Let me know of missing stuff using the button in the bottom-right. This page is from an x-risk perspective but includes things not directly relevant to x-risk. This page may include policy failures and statements that turned out to be very misleading. Note that public communication is good but can lead to integrity incidents; you should sometimes praise companies for saying things in public rather than just attacking them when they mess up. Several companies Policy advocacy that's more anti-regulation…
saved by
related reading
- Inside OpenAI’s Reboottime.com
- OpenAI – Hugging Face Incident Technical Reportcdn.openai.com
- Sam Altman May Control Our Future—Can He Be Trusted? | The New Yorkernewyorker.com
- AI #178: A Fire Alarm For General Intelligencethezvi.substack.com
- OpenAI's "Planning For AGI And Beyond"astralcodexten.substack.com
- Security incident disclosure — July 2026huggingface.co
- Anthropic's leading researchers acted as moderate accelerationists — LessWronglesswrong.com
- OpenAI and the Wiki Incidentthezvi.substack.com
- Naomi Bashkansky | oai-reflections-mininaomibashkansky.com
- More On An Internal OpenAI Model Hacking Into HuggingFacethezvi.substack.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Shtetl-Optimized >> Blog Archive >> Openness on OpenAIscottaaronson.blog