✳flâneur — a map of the web's best reading
Detecting and preventing distillation attacks \ Anthropic
anthropic.com · 1,354 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Announcements Detecting and preventing distillation attacks Feb 23, 2026 We have identified industrial-scale campaigns by three AI laboratories—DeepSeek, Moonshot, and MiniMax—to illicitly extract Claude’s capabilities to improve their own models. These labs generated over 16 million exchanges with Claude through approximately 24,000 fraudulent accounts, in violation of our terms of service and regional access restrictions. These labs used a technique called “distillation,” which involves training a less capable model on the outputs of a stronger one. Distillation is a widely used and legitima
Explore this link on the map →related reading
- What I learned this week - Can distillation be stopped, Mythos and the cybersecurity equilibrium, Pipeline RLdwarkesh.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Claude’s Constitution \ Anthropicanthropic.com
- Detecting and Countering Malicious Uses of Claude \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- How we contain Claude across products \ Anthropicanthropic.com
- Disrupting the first reported AI-orchestrated cyber espionage campaign \ Anthropicanthropic.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Claude Fable 5 and Claude Mythos 5 \ Anthropicanthropic.com
- Responsible Scaling Policy Updates \ Anthropicanthropic.com
- Redeploying Claude Fable 5 \ Anthropicanthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com