flâneur — a map of the web's best reading

Announcing Apollo Research - LessWrong

lesswrong.com · 3,452 words · saved by 1 readers

TL;DR • 1. We are a new AI evals research organization called Apollo Research based in London. 2. We think that strategic AI deception – where a model outwardly seems aligned but is in fact misali…

x Announcing Apollo Research — LessWrong Interpretability (ML & AI) Organization Updates AI Evaluations AI Risk Deceptive Alignment Project Announcement Apollo Research (org) AI Governance AI Personal Blog 226 Announcing Apollo Research by Marius Hobbhahn , beren , Lee Sharkey , Lucius Bushnaq , Dan Braun , Mikita Balesni , Jérémy Scheurer 30th May 2023 AI Alignment Forum 10 min read 11 226 Ω 90 TL;DR We are a new AI evals research organization called Apollo Research based in London. We think that strategic AI deception – where a model outwardly seems aligned but is in fact misaligned – is a c

Explore this link on the map →

related reading