✳flâneur — a map of the web's best reading
Multi-agent safety — AI Alignment Forum
alignmentforum.org · 3,123 words · saved by 1 readers
Note: this post is most explicitly about safety in multi-agent training regimes. However, many of the arguments I make are also more broadly applicab…
x Multi-agent safety — AI Alignment Forum Shaping safer goals AI Frontpage 15 Multi-agent safety by Richard_Ngo 16th May 2020 6 min read 8 15 Note: this post is most explicitly about safety in multi-agent training regimes. However, many of the arguments I make are also more broadly applicable - for example, when training a single agent in a complex environment, challenges arising from the environment could play an analogous role to challenges arising from other agents. In particular, I expect that the diagram in the 'Developing General Intelligence' section will be applicable to most possible
Explore this link on the map →related reading
- Why multi-agent safety is important — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Commentary on AGI Safety from First Principles — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Teaching Claude why \ Anthropicanthropic.com
- ROGUE:arxiv.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- AGI safety career advice — EA Forumforum.effectivealtruism.org
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- Obedient AI - Nina Panicksseryblog.ninapanickssery.com
- Reward Is Not Enough — LessWronglesswrong.com