[2511.09287] From Model Training to Model Raising
Abstract:Current AI training methods align models with human values only after their core capabilities have been established, resulting in models that are easily misaligned and lack deep-rooted value systems. We propose a paradigm shift from "model training" to "model raising", in which alignment is woven into a model's development from the start. We identify several key components for this paradigm, all centered around redesigning the training corpus: reframing training data from a first-person perspective, recontextualizing information as lived experience, simulating social interactions, and scaffolding the ordering of training data. We expect that this redesign of the training corpus will lead to an early commitment to values from the first training token onward, such that knowledge, skills, and values are intrinsically much harder to separate. In an ecosystem in which large language model capabilities start overtaking human capabilities in many tasks, this seems to us like a critical need.
[2511.09287] From Model Training to Model Raising Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Artificial Intelligence arXiv:2511.09287 (cs) [Submitted on 12 Nov 2025 ( v1 ), last revised 17 Nov 2025 (this version, v2)] Title: From Model Training to Model Raising Authors: Roland Aydin , Christian Cyron , Steve Bachelor , Ashton Anderson , Robert West View a PDF of the paper titled From Model Training to Model Raising, by Roland Aydin and 4 other authors View PDF HTML (experiment
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Thomas Larsen's Shortform — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- How far does alignment midtraining generalize?alignment.openai.com
- 1a3orn's Shortform — LessWronglesswrong.com
- Synthetic Persona Pretraining: Alignment from Token Zero — LessWronglesswrong.com
- Alignment faking in large language modelsarxiv.org
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net