What makes a good monitoring prompt? – Apollo Research
We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring (e.g. opposed to real-time binary prompts).
What makes a good monitoring prompt? – Apollo Research July 23, 2026 What makes a good monitoring prompt? Contents Summary We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring (e.g. opposed to real-time binary prompts). Setup: We start with a canonical prompt (~281 lines long) that we split into 15 different components, e.g. definition, reasoning structure, example, etc. We then do various ablations and compressions across five failure modes, ~211–231 trajectories each (1,094 total), spanning the fu
saved by
related reading
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu
- PromptLayer — Prompt Management, Evals & Observabilitypromptlayer.com
- Agentationagentation.dev
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4.github.com
- Why We Are Excited About Confessionsalignment.openai.com
- How to build fast, efficient monitors for AI models using probes - Goodfiregoodfire.com
- Agentationagentation.com
- Early work on monitorability evaluations - METRmetr.org
- Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settingsarxiv.org
- confessions_paper.pdfcdn.openai.com