Measuring Reward-Seeking by Instilling Contrastive Beliefs
We developed Contrastive Synthetic Document Finetuning (Contrastive SDF), a new test for whether an AI model changes its behavior when it has different beliefs about what a grader rewards.
Measuring Reward-Seeking by Instilling Contrastive Beliefs ← Back to OpenAI Alignment Blog Measuring Reward-Seeking by Instilling Contrastive Beliefs Jul 21, 2026 · Axel Højmark (Apollo Research), Jérémy Scheurer (Apollo Research), Jenny Nitishinskaya, Felix Hofstätter (Apollo Research), Jason Wolfe, Theodore Ehrenborg (Apollo Research), Bronson Schoen (Apollo Research), Alexander Meinke (Apollo Research) Correspondence: jenny [at] openai.com Read the paper In Brief We developed a new test, Contrastive Synthetic Document Finetuning (Contrastive SDF), for whether an AI model would change i
saved by
related reading
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- Why do models task game?greaterwrong.com
- [2606.12016] Generalization Hacking: Models Can Game Reinforcement Learning by Preventing Behavioral Generalizationarxiv.org
- [2606.26071] Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignmentarxiv.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- confessions_paper.pdfcdn.openai.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Training a Misaligned Reward Seekeralignment.anthropic.com
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com