Frontier Models are Capable of In-context Scheming — AI Alignment Forum
This is a brief summary of what we believe to be the most important takeaways from our new paper and from our findings shown in the o1 system card. W…
x Frontier Models are Capable of In-context Scheming — AI Alignment Forum AI Evaluations Deceptive Alignment AI Frontpage 89 Frontier Models are Capable of In-context Scheming by Marius Hobbhahn , Alex Meinke , Bronson Schoen , rusheb , Jérémy Scheurer , Mikita Balesni 5th Dec 2024 8 min read 24 89 This is a brief summary of what we believe to be the most important takeaways from our new paper and from our findings shown in the o1 system card. We also specifically clarify what we think we did NOT show. Paper: https://www.apolloresearch.ai/research/scheming-reasoning-evaluations Twitter about p
related reading
- Frontier Models are Capable of In-Context Scheming – Apollo Researchapolloresearch.ai
- How will we update about scheming?blog.redwoodresearch.org
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- Teaching Claude Whyalignment.anthropic.com
- How will we update about scheming? — LessWronglesswrong.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org