AE and AI Alignment
AI development is advancing at an exponential pace. Every leap forward escalates both immense opportunities and (existential) risks. Superficial safety tactics—RLHF, prompt engineering, output filtering—just aren't enough. They're brittle guardrails masking deeper structural misalignments. Recent results have revealed even minimally fine-tuned models capable of producing profoundly harmful outputs, hiding dangerous backdoors, and deceptively faking their own alignment. At AE, our stance is clear and urgent: Alignment isn't solved. It's fundamentally a scientific R&D problem—not merely an engineering challenge—and the stakes of getting this right literally couldn't be higher. AI is rapidly integrating into our minds, our economies, and our militaries—yet we still don't understand how it works. That's already alarming. But when we surveyed top alignment researchers, fewer than one in ten believed today's methods would actually solve the core problem before AGI. That's a crisis. So we'r
AE Studio | AI Alignment Research AE Studio AI Alignment Research - Neglected Approaches to Solving the Alignment Problem Hover over lines to ALIGN them. Then scroll down for more ALIGNMENT! Tap text to ALIGN. Then scroll down for more ALIGNMENT! Alignment is solvable. The real problem? No one's really tried yet. We are, and we're focused where the leverage is highest: the neglected approaches that science forgot. Explore AI Alignment ↓ If you don't give a sh*t, click here → Why Alignment Matters AI development is advancing at an exponential pace. Every leap forward escalates both immense oppo
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- (My understanding of) What Everyone in Technical Alignment is Doing and Why — LessWronglesswrong.com
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- Critical review of Christiano's disagreements with Yudkowsky — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Sequent: Scale and Automation for Higher Confidence in Alignment — Sequentsequent.org
- A minimal viable product for alignment - by Jan Leikealigned.substack.com
- PSA: Almost nobody is directly working on superintelligent alignment — LessWronglesswrong.com