Announcing Apollo Research - LessWrong
TL;DR • 1. We are a new AI evals research organization called Apollo Research based in London. 2. We think that strategic AI deception – where a model outwardly seems aligned but is in fact misali…
x Announcing Apollo Research — LessWrong Interpretability (ML & AI) Organization Updates AI Evaluations AI Risk Deceptive Alignment Project Announcement Apollo Research (org) AI Governance AI Personal Blog 226 Announcing Apollo Research by Marius Hobbhahn , beren , Lee Sharkey , Lucius Bushnaq , Dan Braun , Mikita Balesni , Jérémy Scheurer 30th May 2023 AI Alignment Forum 10 min read 11 226 Ω 90 TL;DR We are a new AI evals research organization called Apollo Research based in London. We think that strategic AI deception – where a model outwardly seems aligned but is in fact misaligned – is a c
Explore this link on the map →related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Off Target | CNAScnas.org
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Spring 2026 Projects - SPARsparai.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Interpretability Will Not Reliably Find Deceptive AI — AI Alignment Forumalignmentforum.org
- [2408.02565] Reasons to Doubt the Impact of AI Risk Evaluationsarxiv.org
- Neel Nanda on Mechanistic Interpretability: Progress, Limits, and Paths to Safer AI — EA Forumforum.effectivealtruism.org
- Research Areas in Benchmark Design and Evaluation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Interpretability Will Not Reliably Find Deceptive AI — LessWronglesswrong.com