flâneur

AI safety is not a model property

normaltech.ai · 2,691 words · saved by 1 readers

Trying to make an AI model that can’t be misused is like trying to make a computer that can’t be used for bad things

The assumption that AI safety is a property of AI models is pervasive in the AI community. It is seen as so obvious that it is hardly ever explicitly stated. Because of this assumption: Companies have made big investments in red teaming their models before releasing them. Researchers are frantically trying to fix the brittleness of model alignment techniques. Some AI safety advocates seek to restrict open models given concerns that they might pose unique risks. Policymakers are trying to find the training compute threshold above which safety risks become serious enough to justify…

related reading