Is AI alignment on track? Is it progressing... too fast? - Alexey Guzey
If you ask an alignment researcher how to measure the alignment of GPT-4 or Claude, they might go on an hour-long tirade about deceptive alignment, instrumental convergence, and the many-worlds interpretation of quantum mechanics — but they won’t give you any numbers. How come? Why do we have AI capabilities benchmarks, “normal” safety/robustness benchmarks, but no alignment benchmark? And how can we measure the alignment of a human, a computer, God or anything else at all? Ok, what about the chances of AI destroying all of humanity or taking over forever, turning us into its slaves or happy but powerless pets (“p-doom”)? When physicists started to worry about the atomic bomb potentially igniting the atmosphere, they did real calculations and got real numbers. But if you ask people about their p-doom, you’ll get numbers anywhere from 0.1% to 99.9%, and there won’t be a rigorous model backing a single one of them. Only stories, narratives, and vague prophecies of doom. Actually, the sto
Is AI alignment on track? Is it progressing... too fast? created: 2023-10-20 Fri ; modified: 2024-06-07 Table of Contents 5 random top posts > Omens of exceptional talent > It Is Your Responsibility to Follow Up > If the moon doesn't need gravity, why do we? The necessity of understanding for general intelligence > Napoleon: a Cautionary Tale for Young Idealists > A Two sentence Jailbreak for GPT-4 and Claude & Why Nobody Knows How to Fix It If you ask an alignment researcher how to measure the alignment of GPT-4 or Claude, they might go on an hour-long tirade about deceptive alignment, instru
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- The state of AI safety in four fake graphs — LessWronglesswrong.com
- gpt-4.pdfcdn.openai.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Alignment remains a hard, unsolved problem — AI Alignment Forumalignmentforum.org
- Why I’m optimistic about our alignment approachaligned.substack.com