Imitative Generalisation (AKA 'Learning the Prior') — AI Alignment Forum
Tl;dr We want to be able to supervise models with superhuman knowledge of the world and how to manipulate it. For this we need an overseer to be able…
x Imitative Generalisation (AKA 'Learning the Prior') — AI Alignment Forum Debate (AI safety technique) Distillation & Pedagogy Iterated Amplification OpenAI Outer Alignment AI Frontpage 60 Imitative Generalisation (AKA 'Learning the Prior') by Beth Barnes 10th Jan 2021 13 min read 15 60 Tl;dr We want to be able to supervise models with superhuman knowledge of the world and how to manipulate it. For this we need an overseer to be able to learn or access all the knowledge our models have, in order to be able to understand the consequences of suggestions or decisions from the model. If the overs
Explore this link on the map →related reading
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Just Ask for Generalization | Eric Jangevjang.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Better priors as a safety problem — LessWronglesswrong.com
- An AI Which Imitates Humans Can Beat Humans | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- [2312.09390] Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervisionar5iv.labs.arxiv.org
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- [1810.08575] Supervising strong learners by amplifying weak expertsar5iv.labs.arxiv.org
- Mediumai-alignment.com