Relaxed adversarial training for inner alignment - AI Alignment Forum
This post is part of research I did at OpenAI with mentoring and guidance from Paul Christiano. It also represents my current agenda regarding what I believe looks like the most promising approach fo…
x Relaxed adversarial training for inner alignment — AI Alignment Forum Inner Alignment AI Risk Interpretability (ML & AI) AI Frontpage 31 Relaxed adversarial training for inner alignment by evhub 10th Sep 2019 33 min read 27 31 This post is part of research I did at OpenAI with mentoring and guidance from Paul Christiano. It also represents my current agenda regarding what I believe looks like the most promising approach for addressing inner alignment . One particularly concerning failure mode for any sort of advanced AI system is for it to have good performance on the training distribution,
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Alignment Faking Mitigationsalignment.anthropic.com
- An overview of 11 proposals for building safe advanced AI — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Reinforcement learning towards broadly and persistently beneficial modelsalignment.openai.com
- How likely is deceptive alignment? — AI Alignment Forumalignmentforum.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org