[2601.21571] Shaping capabilities with token-level data filtering
Abstract:Current approaches to reducing undesired capabilities in language models are largely post hoc, and can thus be easily bypassed by adversaries. A natural alternative is to shape capabilities during pretraining itself. On the proxy task of removing medical capabilities, we show that the simple intervention of filtering pretraining data is highly effective, robust, and inexpensive at scale. Inspired by work on data attribution, we show that filtering tokens is more effective than filtering documents, achieving the same hit to undesired capabilities at a lower cost to benign ones. Training models spanning two orders of magnitude, we then demonstrate that filtering gets more effective with scale: for our largest models, token filtering leads to a 7000x compute slowdown on the forget domain. We also show that models trained with token filtering can still be aligned on the forget domain. Along the way, we introduce a methodology for labeling tokens with sparse autoencoders and distilling cheap, high-quality classifiers. We also demonstrate that filtering can be robust to noisy labels with sufficient pretraining compute.
View PDF HTML (experimental) Abstract:Current approaches to reducing undesired capabilities in language models are largely post hoc, and can thus be easily bypassed by adversaries. A natural alternative is to shape capabilities during pretraining itself. On the proxy task of removing medical capabilities, we show that the simple intervention of filtering pretraining data is highly effective, robust, and inexpensive at scale. Inspired by work on data attribution, we show that filtering tokens is more effective than filtering documents, achieving the same hit to undesired capabilities at a…
saved by
related reading
- [2512.05648] Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMsarxiv.org
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- A Bitter Lesson for Data Filteringarxiv.org
- Modular Pretraining Enables Access Controlalignment.anthropic.com
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- DataRater: Meta-Learned Dataset Curationarxiv.org
- [2509.14786] Pre-training under infinite computearxiv.org
- [2302.08582] Pretraining Language Models with Human Preferencesarxiv.org