✳flâneur — a map of the web's best reading
How we contain Claude across products \ Anthropic
anthropic.com · 4,299 words · saved by 1 readers
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
Twelve months ago, we'd have rejected out of hand the idea of granting Claude access sufficient to take down an internal Anthropic service. Today that level of access is routine, and Anthropic developers are more productive for it. The risk of these deployments has two components: how likely a failure is, and how much damage one could do. Progress on safeguards and model training has steadily driven down the first; the second—the theoretical blast radius—only grows as capabilities and access expand. Yet as agents become capable of doing work that once required a person or even a team, the cost
Explore this link on the map →saved by
related reading
- Scaling Managed Agents: Decoupling the brain from the hands \ Anthropicanthropic.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com
- Claude’s Constitution \ Anthropicanthropic.com
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Responsible Scaling Policy Updates \ Anthropicanthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Measuring AI agent autonomy in practice \ Anthropicanthropic.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- Thoughts on Claude Fable's silent safeguards — LessWronglesswrong.com
- How Claude Code is built - by Gergely Orosznewsletter.pragmaticengineer.com
- What Is Claude? Anthropic Doesn’t Know, Either | The New Yorkernewyorker.com