GPT-6 Astra System Card - OpenAI Deployment Safety Hub
deploymentsafety.openai.com · 8,704 words · saved by 1 readers
Public, mostly-static site to explore OpenAI safety evaluations, system cards, and posts.
Change log September 9, 2026: Updates to the Alignment section: We clarify how our evaluations test alignment generalization, including which were constructed after training and how the honeypot evaluation relates to training and the Hugging Face incident. We also expand on the limitations of this work: the absence of observed failures does not establish reliability across settings and should be interpreted alongside remaining failures, evaluation awareness findings, and monitoring limitations. September 9, 2026: Naming and substance updates to the section on Verbalized Metagaming and…
saved by
related reading
- GPT-6 Astra System Carddeploymentsafety.openai.com
- GPT-6 Astra System Card - OpenAI Deployment Safety Hubdeploymentsafety.openai.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- CAIS AI Dashboarddashboard.safe.ai
- gpt-4.pdfcdn.openai.com
- GPT-5.6 Preview System Card - OpenAI Deployment Safety Hubdeploymentsafety.openai.com
- GPT-5.6 Preview System Carddeploymentsafety.openai.com
- Summary of METR's predeployment evaluation of GPT-5.6 Solmetr.org
- GPT-4openai.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Claude Opus 4.5: Model Card, Alignment and Safetythezvi.substack.com
- Security incident disclosure — July 2026huggingface.co