Imitative Generalisation (AKA 'Learning the Prior') — AI Alignment Forum
Tl;dr We want to be able to supervise models with superhuman knowledge of the world and how to manipulate it. For this we need an overseer to be able…
x Imitative Generalisation (AKA 'Learning the Prior') — AI Alignment Forum Debate (AI safety technique) Distillation & Pedagogy Iterated Amplification OpenAI Outer Alignment AI Frontpage 60 Imitative Generalisation (AKA 'Learning the Prior') by Beth Barnes 10th Jan 2021 13 min read 15 60 Tl;dr We want to be able to supervise models with superhuman knowledge of the world and how to manipulate it. For this we need an overseer to be able to learn or access all the knowledge our models have, in order to be able to understand the consequences of suggestions or decisions from the model. If the overs
related reading
- Just Ask for Generalization | Eric Jangevjang.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Ilya Sutskever — We're moving from the age of scaling to the age of researchdwarkesh.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- Oversight Assistants: Turning Compute into Understandingbounded-regret.ghost.io
- Better priors as a safety problem — LessWronglesswrong.com
- An AI Which Imitates Humans Can Beat Humans | Tom Cunningham – Tom Cunninghamtecunningham.github.io
- [2312.09390] Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervisionar5iv.labs.arxiv.org
- Foundation Models for Oversight | Transluce AItransluce.org
- Unsupervised Elicitationalignment.anthropic.com