flâneur — a map of the web's best reading

Is AI alignment on track? Is it progressing... too fast? - Alexey Guzey

guzey.com · 1,253 words · saved by 1 readers

If you ask an alignment researcher how to measure the alignment of GPT-4 or Claude, they might go on an hour-long tirade about deceptive alignment, instrumental convergence, and the many-worlds interpretation of quantum mechanics — but they won’t give you any numbers. How come? Why do we have AI capabilities benchmarks, “normal” safety/robustness benchmarks, but no alignment benchmark? And how can we measure the alignment of a human, a computer, God or anything else at all? Ok, what about the chances of AI destroying all of humanity or taking over forever, turning us into its slaves or happy but powerless pets (“p-doom”)? When physicists started to worry about the atomic bomb potentially igniting the atmosphere, they did real calculations and got real numbers. But if you ask people about their p-doom, you’ll get numbers anywhere from 0.1% to 99.9%, and there won’t be a rigorous model backing a single one of them. Only stories, narratives, and vague prophecies of doom. Actually, the sto

Is AI alignment on track? Is it progressing... too fast? created: 2023-10-20 Fri ; modified: 2024-06-07 Table of Contents 5 random top posts > Omens of exceptional talent > It Is Your Responsibility to Follow Up > If the moon doesn't need gravity, why do we? The necessity of understanding for general intelligence > Napoleon: a Cautionary Tale for Young Idealists > A Two sentence Jailbreak for GPT-4 and Claude & Why Nobody Knows How to Fix It If you ask an alignment researcher how to measure the alignment of GPT-4 or Claude, they might go on an hour-long tirade about deceptive alignment, instru

Explore this link on the map →

related reading