Stickiness in AI Behavioral Design
Forethought paper on how current AI model specs could shape the behavior of future LLMs by default, and how to spot precedent-setting "wet cement" moments in AI design.
13th May 2026 [AI Narration] Stickiness in AI Behavioral Design Playback speed Volume 0:00 of 25:35 Current model specs aim to shape the behaviors of near-present models, rather than the behaviors of models arbitrarily far into the future. OpenAI writes that their model spec aims to apply “0-3 months ahead of the present.” Anthropic’s Constitution for Claude notes that the document “is likely to change in important ways in the future.” So these documents are presented as provisional guidelines, not as trying to set behavioral standards for the far future. But what if current model…
saved by
related reading
- What Should Go In A Model Spec?forethought.org
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- AI in 2025: gestalt — LessWronglesswrong.com
- My picture of the present in AI — LessWronglesswrong.com
- the void — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Notes on Inference Integritynewsletter.forethought.org
- Introducing the Model Spec | OpenAIopenai.com
- GitHub - suzana-ilic/study_model_behavior: Model Behavior Study Groupgithub.com
- The persona selection model — LessWronglesswrong.com