Introducing the Model Spec | OpenAI
To deepen the public conversation about how AI models should behave, we’re sharing the Model Spec, our approach to shaping desired model behavior. Update on February 12, 2025: We've released an updated version of the Model Spec. This update reinforces our commitments to customizability, transparency, and intellectual freedom to explore, debate, and create with AI without arbitrary restrictions—while ensuring that guardrails remain in place to reduce the risk of real harm. It builds on the foundations we introduced last May, drawing from our experience applying it in varied contexts from alignment research to serving users across the world. You can read more about the update in this blog post. May 8, 2024: We are sharing a first draft of the Model Spec, a new document that specifies how we want our models to behave in the OpenAI API and ChatGPT. We’re doing this because we think it’s important for people to be able to understand and discuss the practical choices involved in shaping mode
May 8, 2024 Safety Research Introducing the Model Spec Model Spec (opens in a new window) Loading… Share Update on February 12, 2025 : We've released an updated version of the Model Spec. This update reinforces our commitments to customizability, transparency, and intellectual freedom to explore, debate, and create with AI without arbitrary restrictions—while ensuring that guardrails remain in place to reduce the risk of real harm. It builds on the foundations we introduced last May, drawing from our experience applying it in varied contexts from alignment research to serving users across the
saved by
related reading
- The Shape of AI | UX Patterns for Artificial Intelligence Designshapeof.ai
- What Should Go In A Model Spec?forethought.org
- GitHub - suzana-ilic/study_model_behavior: Model Behavior Study Groupgithub.com
- Model Spec (2024/05/08)cdn.openai.com
- How confessions can keep language models honest | OpenAIopenai.com
- [2605.24229] How Well Do Models Follow Their Constitutions?arxiv.org
- Uncensored Modelserichartford.com
- Today, we are releasing a research preview of our user model, along with a set of evaluations designed to measure how faithfully user models capture human behavior.persimmon.humansand.ai
- How well do models follow their constitutions? — LessWronglesswrong.com
- Agent behavioragentbehavior.dev
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Model Spec Midtraining: Improving How Alignment Training Generalizesalignment.anthropic.com