flâneur — a map of the web's best reading

How confessions can keep language models honest | OpenAI

openai.com · 2,300 words · saved by 5 readers

We’re sharing an early, proof-of-concept method that trains models to report when they break instructions or take unintended shortcuts. AI systems are becoming more capable, and we want to understand them as deeply as possible—including how and why they arrive at an answer. Sometimes a model takes a shortcut or optimizes for the wrong objective, but its final output still looks correct. If we can surface when that happens, we can better monitor deployed systems, improve training, and increase trust in the outputs. Research by OpenAI and others has shown that AI models can hallucinate⁠, reward-hack, or be dishonest. At the moment, we see the most concerning misbehaviors, such as scheming⁠ (opens in a new window) , only in stress-tests and adversarial evaluations. But as models become more capable and increasingly agentic, even rare forms of misalignment become more consequential, motivating us to invest in methods that help us better detect, understand, and mitigate these risks. This wo

December 3, 2025 Research Publication How confessions can keep language models honest We’re sharing an early, proof-of-concept method that trains models to report when they break instructions or take unintended shortcuts. Read the paper (opens in a new window) Loading… Share AI systems are becoming more capable, and we want to understand them as deeply as possible—including how and why they arrive at an answer. Sometimes a model takes a shortcut or optimizes for the wrong objective, but its final output still looks correct. If we can surface when that happens, we can better monitor deployed sy

Explore this link on the map →

saved by

related reading