✳flâneur — a map of the web's best reading
How far does alignment midtraining generalize?
alignment.openai.com · 2,588 words · saved by 2 readers
Preliminary experiments on alignment and misalignment midtraining, reasoning posttraining, and generalization to chat and agentic evals.
How far does alignment midtraining generalize? ← Back to OpenAI Alignment Blog How far does alignment midtraining generalize? Mar 27, 2026 · Tomek Korbak, Cameron Raymond, Micah Carroll, Marcus Williams, Mikita Balesni, Alan Guo, Jason Wolfe, Akshay Jagadeesh, Ian Kivlichan Today’s AI agents are based on large language models (LLMs) that draw their knowledge and expectations from vast amounts of internet text. This means that AI agents are exposed to AI safety discourse, including discussions of AI misalignment risk and depictions of humans losing control over malicious AI agents, which could
Explore this link on the map →saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Generalization Dynamics of LM Pre-training — Jiaxin Wenjiaxin-wen.github.io
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Teaching Claude why \ Anthropicanthropic.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- Model Spec Midtraining: Improving How Alignment Training Generalizesalignment.anthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Claude is Now Alignment-Pretrained — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Alignment faking in large language modelsarxiv.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org