flâneur — a map of the web's best reading

What makes a good monitoring prompt? – Apollo Research

apolloresearch.ai · 5,791 words · saved by 1 readers

We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring (e.g. opposed to real-time binary prompts).

What makes a good monitoring prompt? – Apollo Research July 23, 2026 What makes a good monitoring prompt? Contents Summary We want to understand which principles work well for high-quality monitoring prompts. We use a 1-10 severity scale and full-trajectory scoring (e.g. opposed to real-time binary prompts). Setup: We start with a canonical prompt (~281 lines long) that we split into 15 different components, e.g. definition, reasoning structure, example, etc. We then do various ablations and compressions across five failure modes, ~211–231 trajectories each (1,094 total), spanning the fu

Explore this link on the map →

related reading