flâneur — a map of the web's best reading

AI safety is not a model property

substack.com · saved by 1 readers

The assumption that AI safety is a property of AI models is pervasive in the AI community. It is seen as so obvious that it is hardly ever explicitly stated. Because of this assumption: Companies have made big investments in red teaming their models before releasing them. Researchers are frantically trying to fix the brittleness of model alignment techniques. Some AI safety advocates seek to restrict open models given concerns that they might pose unique risks. Policymakers are trying to find the training compute threshold above which safety risks become serious enough to justify intervention (and lacking any meaningful basis for picking one, they seem to have converged on 1026 rather arbitrarily). We think these efforts are inherently limited in their effectiveness. That’s because AI safety is not a model property. With a few exceptions, AI safety questions cannot be asked and answered at the levels of models alone. Safety depends to a large extent on the context and the environment i

Explore this link on the map →

saved by