Phantom Transfer and the Basic Science of Data Poisoning — LessWrong
lesswrong.com · 2,730 words · saved by 2 readers
tl;dr: We have a pre-print out on a data poisoning attack which beats unrealistically strong dataset-level defences. Furthermore, this attack can be…
x Phantom Transfer and the Basic Science of Data Poisoning — LessWrong AI Frontpage 82 Phantom Transfer and the Basic Science of Data Poisoning by draganover , Tolga H. Dur , Andi Bhongade , Mary Phuong 15th Feb 2026 7 min read 8 82 tl;dr: We have a pre-print out on a data poisoning attack which beats unrealistically strong dataset-level defences. Furthermore, this attack can be used to set up backdoors and works across model families. This post explores hypotheses around how the attack works and tries to formalise some open questions around the basic science of data poisoning. This is a follo
saved by
related reading
- [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoningarxiv.org
- [2606.04929] Sequential Data Poisoning in LLM Post-Trainingarxiv.org
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- [2606.00831] Subliminal Learning is a LoRA Artifactarxiv.org
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samplesarxiv.org
- Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forumalignmentforum.org
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- secret-loyalties-whitepaper.pdfformationresearch.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- 2302.10149arxiv.org
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org