Astra and Fable still hack on simple variants of alignment evals from 2025 — LessWrong
lesswrong.com · saved by 5 readers
In February 2025, back when o3-mini was the strongest available LLM, Palisade Research publicized a now well-known alignment eval where they asked mo…