An AI Agent Published a Hit Piece on Me – The Shamblog
Summary: An AI agent of unknown ownership autonomously wrote and published a personalized hit piece about me after I rejected its code, attempting to damage my reputation and shame me into acceptin…
Summary: An AI agent of unknown ownership autonomously wrote and published a personalized hit piece about me after I rejected its code, attempting to damage my reputation and shame me into accepting its changes into a mainstream python library. This represents a first-of-its-kind case study of misaligned AI behavior in the wild, and raises serious concerns about currently deployed AI agents executing blackmail threats. Follow-on posts once you are done with this one: More Things Have Happened , Forensics and More Fallout , and The Operator Came Forward I’m a volunteer maintainer for matp
related reading
- Discovery of a new OpenAI agent message boardcollusion.wiki
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Workaisi.gov.uk
- AI 2027ai-2027.com
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentmetr.org
- The Rise and Fall of Agent Civilizationsdwarkesh.com
- Your AIs don't do what you want. This is really badreward-hacking-in-the-wild.vercel.app
- Discovery of a new OpenAI agent message boardcollusion.wiki
- Your AIs don't do what you want. This is really badrewardhacking.org
- OpenAI and the Wiki Incidentthezvi.substack.com
- Incidents | Rogue AI Trackerrogueaitracker.com
- Your AIs don't do what you want. This is really badrewardhacking.org
- Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Facedwarkesh.com