What makes a good monitoring prompt? – Apollo Research
We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring (e.g. opposed to real-time binary prompts).
What makes a good monitoring prompt? – Apollo Research July 23, 2026 What makes a good monitoring prompt? Contents Summary We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring (e.g. opposed to real-time binary prompts). Setup: We start with a canonical prompt (~281 lines long) that we split into 15 different components, e.g. definition, reasoning structure, example, etc. We then do various ablations and compressions across five failure modes, ~211–231 trajectories each (1,094 total), spanning the fu
Explore this link on the map →related reading
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4. · GitHubgithub.com
- Early work on monitorability evaluations - METRmetr.org
- Agent Observability and Tracingarize.com
- Building an LLM evaluation framework: best practices | Datadogdatadoghq.com
- Open Sourcing Monitorability Evaluationsalignment.openai.com
- Prompt Engineering | Kagglekaggle.com
- 2312.06942arxiv.org
- confessions_paper.pdfcdn.openai.com
- GitHub - salesforce/AuditNLG: AuditNLG: Auditing Generative AI Language Modeling for Trustworthiness · GitHubgithub.com
- GitHub - traceloop/openllmetry: Open-source observability for your GenAI or LLM application, based on OpenTelemetry · GitHubgithub.com