Multi-agent safety — AI Alignment Forum
alignmentforum.org · 3,123 words · saved by 1 readers
Note: this post is most explicitly about safety in multi-agent training regimes. However, many of the arguments I make are also more broadly applicab…
x Multi-agent safety — AI Alignment Forum Shaping safer goals AI Frontpage 15 Multi-agent safety by Richard_Ngo 16th May 2020 6 min read 8 15 Note: this post is most explicitly about safety in multi-agent training regimes. However, many of the arguments I make are also more broadly applicable - for example, when training a single agent in a complex environment, challenges arising from the environment could play an analogous role to challenges arising from other agents. In particular, I expect that the diagram in the 'Developing General Intelligence' section will be applicable to most possible
related reading
- Why multi-agent safety is important — LessWronglesswrong.com
- Patterns and problems in multiagent systemsanthropic.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Commentary on AGI Safety from First Principles — AI Alignment Forumalignmentforum.org
- Of Swarms and Sand Godsblog.cosmos-institute.org
- Without specific countermeasures, the easiest path to transformative AI likely leads to AI takeover — LessWronglesswrong.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- ROGUE:arxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- AGI safety career advice — EA Forumforum.effectivealtruism.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com