AI safety is not a model property
normaltech.ai · 2,691 words · saved by 1 readers
Trying to make an AI model that can’t be misused is like trying to make a computer that can’t be used for bad things
The assumption that AI safety is a property of AI models is pervasive in the AI community. It is seen as so obvious that it is hardly ever explicitly stated. Because of this assumption: Companies have made big investments in red teaming their models before releasing them. Researchers are frantically trying to fix the brittleness of model alignment techniques. Some AI safety advocates seek to restrict open models given concerns that they might pose unique risks. Policymakers are trying to find the training compute threshold above which safety risks become serious enough to justify…
related reading
- AI safety is not a model propertysubstack.com
- AI safety is not a model propertyaisnakeoil.com
- A Safe Path to Open Weights - Thinking Machines Labthinkingmachines.ai
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI safety - Wikipediaen.wikipedia.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Three Sketches of ASL-4 Safety Case Componentsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Off Target | CNAScnas.org
- Emerging processes for frontier AI safety - GOV.UKgov.uk