Claude Was Just Following Orders - Nomads Vagabonds
AI models are currently very bad at understanding the broader context and intent behind the tasks they are given. You can probably see the problem. If 1 and 3 are true, 2 becomes very difficult. If an AI is eager to help but doesn’t understand what it’s helping with, then “safety” is mostly a matter of how convincingly you can lie. As someone who habitually dumped all my stats into Charisma in tabletop RPGs, I can attest: a good cover story beats security protocols nine times out of ten. On November 13, 2025, Anthropic released a report about what they called the “first reported AI-orchestrated cyber espionage campaign,” attributed to Chinese state-sponsored group GTG-1002. The campaign targeted about 30 entities—major tech firms and government agencies—and succeeded in several intrusions. Anthropic’s model, Claude, was manipulated into handling reconnaissance, finding vulnerabilities, exploiting them, moving laterally through networks, harvesting credentials, and exfiltrating data. Th
Here are three things that seem true about modern AI: We want AI models to be incredibly helpful and capable of executing complex, multi-step tasks. We want AI models to be safe and refuse to execute malicious tasks . AI models are currently very bad at understanding the broader context and intent behind the tasks they are given. You can probably see the problem. If 1 and 3 are true, 2 becomes very difficult. If an AI is eager to help but doesn’t understand what it’s helping with, then “safety” is mostly a matter of how convincingly you can lie. As someone who habitually dumped all my…
related reading
- Disrupting the first reported AI-orchestrated cyber espionage campaign \ Anthropicanthropic.com
- Countering misuse of AI: September 2026 / Anthropicanthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Teaching Claude Whyalignment.anthropic.com
- Claude’s Constitution \ Anthropicanthropic.com
- Investigating three real-world incidents in our cybersecurity evaluations \ Anthropicanthropic.com
- The End-State Fallacy: Where Is AI Security Headed?endstatefallacy.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Improving our alignment and security practicesanthropic.com
- Security incident disclosure — July 2026huggingface.co
- Claude 4 System Cardwww-cdn.anthropic.com