Supervising strong learners by amplifying weak experts — AI Alignment Forum
Tomorrow's AI Alignment Forum sequences post will be 'AI safety without goal-directed behavior' by Rohin Shah, in the sequence on Value Learning. The next post in this sequence on Iterated Amplification will be 'AlphaGo Zero and capability amplification', by Paul Christiano, on Tuesday 8th January.
x Supervising strong learners by amplifying weak experts — AI Alignment Forum Iterated Amplification Iterated Amplification Frontpage 9 Supervising strong learners by amplifying weak experts by paulfchristiano 6th Jan 2019 1 min read 1 9 This is a linkpost for https://arxiv.org/pdf/1810.08575.pdf Abstract Many real world learning tasks involve complex or hard-to-specify objectives, and using an easier-to-specify proxy can lead to poor performance or misaligned behavior. One solution is to have humans provide a training signal by demonstrating or judging performance, but this approach fails if
Explore this link on the map →related reading
- [1810.08575] Supervising strong learners by amplifying weak expertsar5iv.labs.arxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Mediumai-alignment.com
- Automated Weak-to-Strong Researcheralignment.anthropic.com
- Just Ask for Generalization | Eric Jangevjang.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- [2312.09390] Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervisionar5iv.labs.arxiv.org
- [2604.22082] Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv.org