Measuring Reward-Seeking by Instilling Contrastive Beliefs
We developed Contrastive Synthetic Document Finetuning (Contrastive SDF), a new test for whether an AI model changes its behavior when it has different beliefs about what a grader rewards.
Measuring Reward-Seeking by Instilling Contrastive Beliefs ← Back to OpenAI Alignment Blog Measuring Reward-Seeking by Instilling Contrastive Beliefs Jul 21, 2026 · Axel Højmark (Apollo Research), Jérémy Scheurer (Apollo Research), Jenny Nitishinskaya, Felix Hofstätter (Apollo Research), Jason Wolfe, Theodore Ehrenborg (Apollo Research), Bronson Schoen (Apollo Research), Alexander Meinke (Apollo Research) Correspondence: jenny [at] openai.com Read the paper In Brief We developed a new test, Contrastive Synthetic Document Finetuning (Contrastive SDF), for whether an AI model would change i
Explore this link on the map →related reading
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Models Don't "Get Reward" — LessWronglesswrong.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Training on Documents about Reward Hacking Induces Reward Hackingalignment.anthropic.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com