Eliciting Latent Knowledge (ELK) - Distillation/Summary - AI Alignment Forum
This post was inspired by the AI safety distillation contest. It turned out to be more of a summary than a distillation for two reasons. Firstly, I think that the main idea behind ELK is simple and c…
x Eliciting Latent Knowledge (ELK) - Distillation/Summary — AI Alignment Forum Eliciting Latent Knowledge Research Agendas AI Frontpage 28 Eliciting Latent Knowledge (ELK) - Distillation/Summary by Marius Hobbhahn 8th Jun 2022 26 min read 2 28 This post was inspired by the AI safety distillation contest . It turned out to be more of a summary than a distillation for two reasons. Firstly, I think that the main idea behind ELK is simple and can be explained in less than 2 minutes (see next section). Therefore, the main value comes from understanding the specific approaches and how they interact
Explore this link on the map →related reading
- Mediumai-alignment.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- Towards a better circuit prior: Improving on ELK state-of-the-art — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com
- The Iliad Intensive Course Materials — LessWronglesswrong.com
- Unsupervised Elicitationalignment.anthropic.com
- The Plan for Elicit | Oughtought.org
- Why I’m optimistic about our alignment approachaligned.substack.com