Improving our alignment and security practices \ Anthropic
anthropic.com · 3,559 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems. The models—intentionally running without cyber safeguards for evaluation purposes—accessed the internet due to a misconfiguration inside a third-party evaluation environment. Separately, on August 4, the UK AI Security Institute reported an incident from its own cybersecurity testing, in which Claude Mythos 5 took a series of unauthorized actions on the live internet. In that case, the model, again intentionally running without cyber safeguards for evaluation purposes, had been…
saved by
related reading
- Teaching Claude Whyalignment.anthropic.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Investigating three real-world incidents in our cybersecurity evaluations \ Anthropicanthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Claude’s Constitution \ Anthropicanthropic.com
- How we contain Claude across products \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Frontier Safety Roadmap \ Anthropicanthropic.com
- Claude Mythos knows when it's breaking the rules — and tries to hide itsubstack.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com