flâneur — a map of the web's best reading

Relaxed adversarial training for inner alignment - AI Alignment Forum

alignmentforum.org · 10,663 words · saved by 1 readers

This post is part of research I did at OpenAI with mentoring and guidance from Paul Christiano. It also represents my current agenda regarding what I believe looks like the most promising approach fo…

x Relaxed adversarial training for inner alignment — AI Alignment Forum Inner Alignment AI Risk Interpretability (ML & AI) AI Frontpage 31 Relaxed adversarial training for inner alignment by evhub 10th Sep 2019 33 min read 27 31 This post is part of research I did at OpenAI with mentoring and guidance from Paul Christiano. It also represents my current agenda regarding what I believe looks like the most promising approach for addressing inner alignment . One particularly concerning failure mode for any sort of advanced AI system is for it to have good performance on the training distribution,

Explore this link on the map →

related reading