flâneur

GPT-6 Astra System Card - OpenAI Deployment Safety Hub

deploymentsafety.openai.com · 8,704 words · saved by 1 readers

Today, we are releasing GPT-6 Astra, the most capable model we have ever broadly deployed. Astra is our first model to reach the Critical level of cybersecurity capability under our Preparedness Framework.

Change log September 9, 2026: Updates to the Alignment section: We clarify how our evaluations test alignment generalization, including which were constructed after training and how the honeypot evaluation relates to training and the Hugging Face incident. We also expand on the limitations of this work: the absence of observed failures does not establish reliability across settings and should be interpreted alongside remaining failures, evaluation awareness findings, and monitoring limitations. September 9, 2026: Naming and substance updates to the section on Verbalized Metagaming and…

saved by

related reading