✳flâneur — a map of the web's best reading
Phantom Transfer and the Basic Science of Data Poisoning — LessWrong
lesswrong.com · 2,730 words · saved by 1 readers
tl;dr: We have a pre-print out on a data poisoning attack which beats unrealistically strong dataset-level defences. Furthermore, this attack can be…
x Phantom Transfer and the Basic Science of Data Poisoning — LessWrong AI Frontpage 82 Phantom Transfer and the Basic Science of Data Poisoning by draganover , Tolga H. Dur , Andi Bhongade , Mary Phuong 15th Feb 2026 7 min read 8 82 tl;dr: We have a pre-print out on a data poisoning attack which beats unrealistically strong dataset-level defences. Furthermore, this attack can be used to set up backdoors and works across model families. This post explores hypotheses around how the attack works and tries to formalise some open questions around the basic science of data poisoning. This is a follo
Explore this link on the map →saved by
related reading
- [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoningarxiv.org
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samplesarxiv.org
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Probe-Based Data Attribution: Surfacing and Mitigating Undesirable Behaviors in LLM Post-Traininggoodfire.ai
- How do LLMs generalize when we do training that is intuitively compatible with two off-distribution behaviors? — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- secret-loyalties-whitepaper.pdfformationresearch.com
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com