Richard Ngo's Shortform — LessWrong
Comment by Richard_Ngo - In response to an email about what a pro-human ideology for the future looks like, I wrote up the following: The pro-human egregore I'm currently designing (which I call fractal empowerment) incorporates three key ideas: Firstly, we can see virtue ethics as a way for less powerful agents to aggregate to form more powerful superagents that preserve the interests of those original less powerful agents. E.g. virtues like integrity, loyalty, etc help prevent divide-and-conquer strategies. This would have been in the interests of the rest of the world when Europe was trying to colonize them, and will be in the best interests of humans when AIs try to conquer us. Secondly, the most robust way for a more powerful agent to be altruistic towards a less powerful agent is not for it to optimize for that agent's welfare, but rather to optimize for its empowerment. This prevents predatory strategies from masquerading as altruism (e.g. agents claiming "I'll conquer you and then I'll empower you", which then somehow never get around to the second step). Thirdly: the generational contract. From any given starting point, there are a huge number of possible coalitions which could form, and in some sense it's arbitrary which set of coalitions you choose. But one thing which is true for both humans and AIs is that each generation wants to be treated well by the next generation. And so the best intertemporal Schelling point is for coalitions to be inherently historical: that is, they balance the interests of old agents and new agents (even when the new agents could in theory form a coalition against all the old agents). From this perspective, path-dependence is a feature not a bug: there are many possible futures but only one history, meaning that this single history can be used to coordinate. In some sense this is a core idea of UDT: when coordinating with forks of yourself, you defer to your unique last common ancestor. When it's not literally a fork of yourself, there's more arb
x Richard Ngo's Shortform — LessWrong Richard Ngo's Shortform by Richard_Ngo 26th Apr 2020 AI Alignment Forum 1 min read 791 6 Ω 3 This is a special post for quick takes by Richard_Ngo . Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page . Rendering 0 / 791 comments, sorted by top scoring (show more) Click to highlight new comments since: Today at 2:08 PM Moderation Log More from Richard_Ngo View more Curated and popular this week 791 Comments 791 Comment Permalink Richard_Ngo 1y * 97 22 In response to an email about what a pro-human i
Explore this link on the map →related reading
- ⿻ Symbiogenesis vs. Convergent Consequentialism — LessWronglesswrong.com
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- Gradual Paths to Collective Flourishing — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Solipsistic Superintelligence is Unlikely to be Cooperativearxiv.org
- Making deals with early schemers — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- The Risk of Gradual Disempowerment from AIthezvi.substack.com
- We're already in AI takeoff — LessWronglesswrong.com
- The Best of LessWrong — LessWronglesswrong.com
- Hyperstition for Good: Write the Future All Sentient Beings Deservehyperstition.sentientfutures.ai
- Why AIs aren't power-seeking yet — LessWronglesswrong.com