Frontier Models are Capable of In-context Scheming — AI Alignment Forum
This is a brief summary of what we believe to be the most important takeaways from our new paper and from our findings shown in the o1 system card. W…
x Frontier Models are Capable of In-context Scheming — AI Alignment Forum AI Evaluations Deceptive Alignment AI Frontpage 89 Frontier Models are Capable of In-context Scheming by Marius Hobbhahn , Alex Meinke , Bronson Schoen , rusheb , Jérémy Scheurer , Mikita Balesni 5th Dec 2024 8 min read 24 89 This is a brief summary of what we believe to be the most important takeaways from our new paper and from our findings shown in the o1 system card. We also specifically clarify what we think we did NOT show. Paper: https://www.apolloresearch.ai/research/scheming-reasoning-evaluations Twitter about p
Explore this link on the map →related reading
- Frontier Models are Capable of In-Context Scheming – Apollo Researchapolloresearch.ai
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- How will we update about scheming? — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com