✳flâneur — a map of the web's best reading
Goodfire on X: "Have you debugged your training data? You might not like what you find. Introducing predictive data debugging: reveal and shape what your model will learn before training. In DPO datasets, we found broken guardrails, hallucinations, and fish fart fan fiction (seriously). (1/9) https://t.co/V9MrEQvBvq" / X
x.com · 43 words · saved by 1 readers
Have you debugged your training data? You might not like what you find. Introducing predictive data debugging: reveal and shape what your model will learn before training. In DPO datasets, we found broken guardrails, hallucinations, and fish fart fan fiction (seriously). (1/9) https://t.co/V9MrEQvBvq
@GoodfireAI: Have you debugged your training data? You might not like what you find. Introducing predictive data debugging: reveal and shape what your model will learn before training. In DPO datasets, we found broken guardrails, hallucinations, and fish fart fan fiction (seriously). (1/9)
Explore this link on the map →saved by
related reading
- [2606.12360] Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signalarxiv.org
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- RL Pet Peeves Part 1 · Aurielaurielws.github.io
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- [2602.05910] Chunky Post-Training: Data Driven Failures of Generalizationarxiv.org
- Why you need to improve your training data, and how to do it << Pete Warden's blogpetewarden.com
- [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoningarxiv.org
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Prompt Engineering | Kagglekaggle.com
- Trainloop AItrainloop.ai
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org