Legible vs. Illegible AI Safety Problems — LessWrong
Some AI safety problems are legible (obvious or understandable) to company leaders and government policymakers, implying they are unlikely to deploy or allow deployment of an AI while those problems remain open (i.e., appear unsolved according to the information they have access to). But some problems are illegible (obscure or hard to understand, or in a common cognitive blind spot), meaning there is a high risk that leaders and policymakers will decide to deploy or allow deployment even if they are not solved. (Of course, this is a spectrum, but I am simplifying it to a binary for ease of exposition.) From an x-risk perspective, working on highly legible safety problems has low or even negative expected value. Similar to working on AI capabilities, it brings forward the date by which AGI/ASI will be deployed, leaving less time to solve the illegible x-safety problems. In contrast, working on the illegible problems (including by trying to make them more legible) does not have this issu
x Legible vs. Illegible AI Safety Problems — LessWrong AI Risk AI Curated 2025 Top Fifty: 49 % 396 Legible vs. Illegible AI Safety Problems by Wei Dai 4th Nov 2025 AI Alignment Forum 2 min read 96 396 Ω 111 Some AI safety problems are legible (obvious or understandable) to company leaders and government policymakers, implying they are unlikely to deploy or allow deployment of an AI while those problems remain open (i.e., appear unsolved according to the information they have access to). But some problems are illegible (obscure or hard to understand, or in a common cognitive blind spot), meanin
Explore this link on the map →related reading
- Legible vs. Illegible AI Safety Problems — EA Forumforum.effectivealtruism.org
- Wei Dai — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Dario Amodei — The Urgency of Interpretabilitydarioamodei.com
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- The Legibility Problempress.asimov.com
- Limits to Legibility — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- AI Safety Seems Hard to Measurecold-takes.com
- AI Safety | Arkosevictoriabrook.github.io
- AI Safety for Fleshy Humans: a whirlwind touraisafety.dance
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org