Wei Dai's Shortform — LessWrong
Comment by Wei Dai - Long horizon agency / strategic competence approximately does not exist among humans, even the smartest ones. With very few exceptions, billionaires spend or give away their money haphazardly, philosophers don't bother to think about long term implications of AI on philosophy production (positive or negative), Terence Tao spends his time wireheading on abstract math instead of doing anything remotely like instrumental convergence. Unlike my youthful expectations (upon reading Vernor Vinge), there are no university departments filled with super-geniuses charting a path for humanity to safely navigate the Singularity. Aside from this, humans also have a bunch of other safety problems, like being bad at philosophy, being easy to manipulate, having strange and unstable values, tending to ignore risks they create (because acknowledging them would be bad for one's status). So if you try to improve people's agency, you likely just end up getting people like founders of FTX and OAI. What about getting help from AI? Well they seem to suffer from many of the same safety problems, but in even more severe forms. E.g., current AI capabilities are even more skewed towards short-horizon, easily verifiable tasks, like math and coding. They seem even more prone to reward gaming, are even worse at doing philosophy, are liable to have even more alien values, etc. Both AI and human safety seem to have this interlocking nature, i.e., there is a bunch of different safety problems where solving some but not all of them at the same time can make the overall situation worse. (For example, solving AI intent alignment allows humanity to do more damage to itself with AI help, if AI doesn't also provide competent strategic and philosophical assistance, but increasing AI strategic competence risks allowing misaligned AI to take over more easily.) This feature demands a high level of strategic competence to recognize and navigate, which is just what we don't have. I've been supportive of AI p
x Wei Dai's Shortform — LessWrong Wei Dai's Shortform by Wei Dai 1st Mar 2024 AI Alignment Forum 1 min read 595 10 Ω 6 This is a special post for quick takes by Wei Dai . Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page . Rendering 0 / 595 comments, sorted by top scoring (show more) Click to highlight new comments since: Today at 1:58 PM Moderation Log More from Wei Dai View more Curated and popular this week 595 Comments 595 Comment Permalink Wei Dai 5mo 142 8 Long horizon agency / strategic competence approximately does not exist a
Explore this link on the map →related reading
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- You will be OK — LessWronglesswrong.com
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- AI will change the world, but won’t take it over by playing “3-dimensional chess”. – Windows On Theorywindowsontheory.org
- Wei Dai's Shortform — AI Alignment Forumalignmentforum.org
- My AI Opinions - by Scott Alexander - Astral Codex Tensubstack.com
- AI Safety for Fleshy Humans: a whirlwind touraisafety.dance
- Humans are not automatically strategic — LessWronglesswrong.com
- Post 50: Good Research Takes are Not Sufficient for Good Strategic Takes - Neel Nandaneelnanda.io
- 2023 - by Dean W. Ball - Hyperdimensionalhyperdimensional.co
- The Best of LessWrong — LessWronglesswrong.com
- Clarifying and predicting AGI — LessWronglesswrong.com