flâneur — a map of the web's best reading

Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training

goodfire.ai · 5,450 words · saved by 1 readers

Post-training can introduce undesired side effects that are difficult to detect and even harder to trace to specific training datapoints. We show that a probe-based method can surface concerning behaviors that emerge during LLM post-training, and that probes can identify the datapoints responsible for a specific harmful behavior. Filtering out those datapoints and retraining significantly reduces the behavior. We introduce a natural testbed for data attribution: a harmful behavior that emerges during DPO training of OLMo 2 7B, where the model learns to comply with certain harmful requests that it previously refused. Filtering out the data flagged by our probe reduces the harmful behavior by 63% without compromising general performance — outperforming gradient-based and LLM-judge alternatives at one tenth the cost once the probe is trained. Swapping the accepted and rejected labels on the same datapoints reduces the behavior by 78%, suggesting the preference data contains systematically

Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training Research Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Training Authors Frank Xiao * Santiago Aranguri † * California Institute of Technology (work done as a SPAR mentee) † Goodfire Published April 29, 2026 Full Paper Read on arXiv → Contents Introduction 1. A Naturally Occurring Harmful Behavior in DPO 2. Probe-Based Data Attribution Identifying Problematic Datapoints Mitigating a Harmful Behavior via Dataset Intervention Filtering datapoints Swapping la

Explore this link on the map →

saved by

related reading