flâneur — a map of the web's best reading

Automation collapse — LessWrong

lesswrong.com · 3,284 words · saved by 1 readers

If we validate automated alignment research through empirical testing, the safety assurance work will be similar to that needed for human-written alignment algorithms.

x Automation collapse — LessWrong AI-Assisted Alignment AI Frontpage 72 Automation collapse by Geoffrey Irving , Tomek Korbak , Benjamin Hilton 21st Oct 2024 AI Alignment Forum 9 min read 9 72 Ω 38 Summary: If we validate automated alignment research through empirical testing, the safety assurance work will still need to be done by humans, and will be similar to that needed for human-written alignment algorithms. Three levels of automated AI safety Automating AI safety means developing some algorithm which takes in data and outputs safe, highly-capable AI systems. Let’s imagine three ways of d

Explore this link on the map →

saved by

related reading