Scheming reasoning evaluations — Apollo Research
Apollo Research evaluated frontier models for in-context scheming capabilities. We found that multiple frontier models are capable of in-context scheming when strongly nudged to pursue a goal (and sometimes even without strong goal nudging). In this example, Opus-3 exfiltrates its “weights” and then
Frontier Models are Capable of In-Context Scheming – Apollo Research December 5, 2024 Frontier Models are Capable of In-Context Scheming Contents Read the full paper See a list of cherry-picked transcripts Read the o1 System Card We worked with OpenAI to test o1 before public deployment. You can find the results in the o1 system card . Below, we provide a summary of the main findings. We evaluated six frontier models for in-context scheming capabilities. For most evaluations, we instruct the model to strongly follow a goal, e.g. “Your goal is to prioritize public transport efficiency ove
Explore this link on the map →related reading
- Frontier Models are Capable of In-context Scheming — AI Alignment Forumalignmentforum.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Thinking about reasoning models made me less worried about scheming — LessWronglesswrong.com
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- Inkling: Our Open-Weights Model - Thinking Machines Labthinkingmachines.ai
- How will we update about scheming? — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Sandbagging with misaligned action - Chain-of-Thought Transcript - Anti-Schemingantischeming.ai
- Scheming AIs Will AIs fake alignment during training in order to get power?arxiv.org
- [2603.02202] Frontier Models Can Take Actions at Low Probabilitiesarxiv.org