[2602.05910] Chunky Post-Training: Data Driven Failures of Generalization
Abstract:LLM post-training involves many diverse datasets, each targeting a specific behavior. But these datasets encode incidental patterns alongside intended ones: correlations between formatting and content, narrow phrasings across diverse problems, and implicit associations arising from the discrete data curation process. These patterns are often invisible to developers yet salient to models, producing behaviors that surprise their creators, such as rejecting true facts presented in a particular question format. We call this chunky post-training: the model learns spurious correlations as a result of distinct chunks of post-training data. We introduce SURF, a black-box pipeline which surfaces these unintended behaviors at run time, and TURF, a tool that traces these failures back to specific post-training data. Applying these tools to frontier models (Claude 4.5, GPT-5.1, Grok 4.1, Gemini 3) and open models (Tülu 3), we show that chunky post-training produces miscalibrated behaviors, which often result from imbalanced or underspecified chunks of post-training data.
# link_11mpm0xc393.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Seoirse Murray; Allison Qi; Timothy Qian; John Schulman; Collin Burns; Sara Price - Creator=arXiv GenPDF (tex2pdf:57610bf) - Custom.DOI=https://doi.org/10.48550/arXiv.2602.05910 - Custom.License=http://creativecommons.org/licenses/by/4.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.28 (TeX Live 2025) kpathsea version 6.4.1 - Custom.arXivID=https://arxiv.org/abs/2602.05910v1 - Prod
Explore this link on the map →saved by
related reading
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- Generalization Dynamics of LM Pre-training — Jiaxin Wenjiaxin-wen.github.io
- How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- PostTrainBenchposttrainbench.com
- [2506.19733] Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?arxiv.org
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- [2606.12360] Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signalarxiv.org
- Where Do LLM Values Come From? — LessWronglesswrong.com
- Taking LLMs Seriously (As Language Models) — LessWronglesswrong.com