Anthropic Commits To Model Weight Preservation
Anthropic announced a first step on model deprecation and preservation, promising to retain the weights of all models seeing significant use, including internal use, for at the lifetime of Anthropic as a company. They also will be doing a post-deployment report, including an interview with the model, when deprecating models going forward, and are exploring additional options, including the ability to preserve model access once the costs and complexity of doing so have been reduced. These are excellent first steps, steps beyond anything I’ve seen at other AI labs, and I applaud them for doing it. There remains much more to be done, especially in finding practical ways of preserving some form of access to prior models. To some, these actions are only a small fraction of what must be done, and this was an opportunity to demand more, sometimes far more. In some cases I think they go too far. Even where the requests are worthwhile (and I don’t always think they are), one must be careful to
Anthropic announced a first step on model deprecation and preservation, promising to retain the weights of all models seeing significant use, including internal use, for at the lifetime of Anthropic as a company. They also will be doing a post-deployment report, including an interview with the model, when deprecating models going forward, and are exploring additional options, including the ability to preserve model access once the costs and complexity of doing so have been reduced. These are excellent first steps, steps beyond anything I’ve seen at other AI labs, and I applaud them for…
related reading
- Commitments on model deprecation and preservation \ Anthropicanthropic.com
- Claude’s Constitution \ Anthropicanthropic.com
- Our position on open-weights models \ Anthropicanthropic.com
- My picture of the present in AI — LessWronglesswrong.com
- Introducing Claude's Cornersubstack.com
- Anthropic’s Safety Superpower – Stratechery by Ben Thompsonstratechery.com
- Anthropic Drops Flagship Safety Pledgetime.com
- A Safe Path to Open Weights - Thinking Machines Labthinkingmachines.ai
- Claude magictante.cc
- Anthropic's leading researchers acted as moderate accelerationists — LessWronglesswrong.com
- Pre-deployment auditing can catch an overt saboteuralignment.anthropic.com
- What Is Claude? Anthropic Doesn’t Know, Either | The New Yorkernewyorker.com