Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training
Post-training can introduce undesired side effects that are difficult to detect and even harder to trace to specific training datapoints. We show that a probe-based method can surface concerning behaviors that emerge during LLM post-training, and that probes can identify the datapoints responsible for a specific harmful behavior. Filtering out those datapoints and retraining significantly reduces the behavior. We introduce a natural testbed for data attribution: a harmful behavior that emerges during DPO training of OLMo 2 7B, where the model learns to comply with certain harmful requests that it previously refused. Filtering out the data flagged by our probe reduces the harmful behavior by 63% without compromising general performance — outperforming gradient-based and LLM-judge alternatives at one tenth the cost once the probe is trained. Swapping the accepted and rejected labels on the same datapoints reduces the behavior by 78%, suggesting the preference data contains systematically
Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training Research Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training Authors Frank Xiao * Santiago Aranguri † * California Institute of Technology (work done as a SPAR mentee) † Goodfire Published April 29, 2026 Full Paper Read on arXiv → Contents Introduction 1. A Naturally Occurring Harmful Behavior in DPO 2. Probe-Based Data Attribution Identifying Problematic Datapoints Mitigating a Harmful Behavior via Dataset Intervention Filtering datapoints Swapping la
Explore this link on the map →saved by
related reading
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Where Do LLM Values Come From? — LessWronglesswrong.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- Surfacing Pathological Behaviors in Language Models | Transluce AItransluce.org
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org