Detecting and preventing distillation attacks \ Anthropic
anthropic.com · 1,354 words · saved by 5 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Announcements Detecting and preventing distillation attacks Feb 23, 2026 We have identified industrial-scale campaigns by three AI laboratories—DeepSeek, Moonshot, and MiniMax—to illicitly extract Claude’s capabilities to improve their own models. These labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts, in violation of our terms of service and regional access restrictions. These labs used a technique called “distillation,” which involves training a less capable model on the outputs of a stronger one. Distillation is a widely used and legitima
saved by
related reading
- Countering misuse of AI: September 2026 / Anthropicanthropic.com
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- Claude’s Constitution \ Anthropicanthropic.com
- The Claude Code Source Leak: fake tools, frustration regexes, undercover mode, and more | Alex Kim's blogalex000kim.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- How to Buy Cheap Claude Tokens in Chinachinatalk.media
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- The Myth of unsafe Open Source AIflorianbrand.com
- Claude Fable 5 and Claude Mythos 5 \ Anthropicanthropic.com
- Stolen Thoughtsstolen-thoughts.com
- How (some) Chinese AI Practitioners View Model Distillationgeopolitechs.org
- Incriminating misaligned AI models via distillation — LessWronglesswrong.com