Strategy For Conditioning Generative Models - LessWrong
This post was written under the mentorship of Evan Hubinger, and assisted by discussions with Adam Jermyn and Johannes Treutlein. See also their previous posts on this project. …
x Strategy For Conditioning Generative Models — LessWrong Language Models (LLMs) Oracle AI AI Frontpage 31 Strategy For Conditioning Generative Models by james.lucassen , evhub 1st Sep 2022 AI Alignment Forum 22 min read 4 31 Ω 16 This post was written under the mentorship of Evan Hubinger, and assisted by discussions with Adam Jermyn and Johannes Treutlein. See also their previous posts on this project. Summary Conditioning Generative Models (CGM) is a strategy to accelerate alignment research using powerful semi-aligned future language models, which seems potentially promising but constraine
Explore this link on the map →related reading
- Cyborgism — LessWronglesswrong.com
- Simulators — LessWronglesswrong.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Teaching Claude why \ Anthropicanthropic.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Cyborgism — AI Alignment Forumalignmentforum.org
- You can, in fact, bamboozle an unaligned AI into sparing your life — LessWronglesswrong.com
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com
- Predicting model behavior before release by simulating deployment | OpenAIopenai.com