Introducing the Model Spec | OpenAI
To deepen the public conversation about how AI models should behave, we’re sharing the Model Spec, our approach to shaping desired model behavior. Update on February 12, 2025: We've released an updated version of the Model Spec. This update reinforces our commitments to customizability, transparency, and intellectual freedom to explore, debate, and create with AI without arbitrary restrictions—while ensuring that guardrails remain in place to reduce the risk of real harm. It builds on the foundations we introduced last May, drawing from our experience applying it in varied contexts from alignment research to serving users across the world. You can read more about the update in this blog post. May 8, 2024: We are sharing a first draft of the Model Spec, a new document that specifies how we want our models to behave in the OpenAI API and ChatGPT. We’re doing this because we think it’s important for people to be able to understand and discuss the practical choices involved in shaping mode
May 8, 2024 Safety Research Introducing the Model Spec Model Spec (opens in a new window) Loading… Share Update on February 12, 2025 : We've released an updated version of the Model Spec. This update reinforces our commitments to customizability, transparency, and intellectual freedom to explore, debate, and create with AI without arbitrary restrictions—while ensuring that guardrails remain in place to reduce the risk of real harm. It builds on the foundations we introduced last May, drawing from our experience applying it in varied contexts from alignment research to serving users across the
Explore this link on the map →saved by
related reading
- Model Spec (2024/05/08)cdn.openai.com
- How confessions can keep language models honest | OpenAIopenai.com
- The Shape of AI | UX Patterns for Artificial Intelligence Designshapeof.ai
- How well do models follow their constitutions? — LessWronglesswrong.com
- Thoughts on Claude’s Constitution – Windows On Theorywindowsontheory.org
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- (a) Anthropic constitution.arxiv.org
- Stress-testing model specs reveals character differences among language modelsalignment.anthropic.com
- GitHub - suzana-ilic/study_model_behavior: Model Behavior Study Group · GitHubgithub.com
- confessions_paper.pdfcdn.openai.com
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Tragithub.com