Goodfire on X: "Have you debugged your training data? You might not like what you find. Introducing predictive data debugging: reveal and shape what your model will learn before training. In DPO datasets, we found broken guardrails, hallucinations, and fish fart fan fiction (seriously). (1/9) https://t.co/V9MrEQvBvq" / X
x.com · 43 words · saved by 1 readers
Have you debugged your training data? You might not like what you find. Introducing predictive data debugging: reveal and shape what your model will learn before training. In DPO datasets, we found broken guardrails, hallucinations, and fish fart fan fiction (seriously). (1/9) https://t.co/V9MrEQvBvq
@GoodfireAI: Have you debugged your training data? You might not like what you find. Introducing predictive data debugging: reveal and shape what your model will learn before training. In DPO datasets, we found broken guardrails, hallucinations, and fish fart fan fiction (seriously). (1/9)
saved by
related reading
- [2606.12360] Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signalarxiv.org
- Dylan Samdsam99.github.io
- Goodfire AIgoodfire.ai
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Datacurve | The data engine for frontier AIdatacurve.ai
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- RL Pet Peeves Part 1 · Aurielaurielws.github.io
- Why you need to improve your training data, and how to do it << Pete Warden's blogpetewarden.com
- Expert Data for Frontier AI - AfterQueryafterquery.com
- Foundation Models for Oversight | Transluce AItransluce.org