✳flâneur — a map of the web's best reading
A small number of samples can poison LLMs of any size \ Anthropic
anthropic.com · 2,058 words · saved by 5 readers
Anthropic research on data-poisoning attacks in large language models
Alignment A small number of samples can poison LLMs of any size Oct 9, 2025 Read the paper In a joint study with the UK AI Security Institute and the Alan Turing Institute, we found that as few as 250 malicious documents can produce a "backdoor" vulnerability in a large language model—regardless of model size or training data volume. Although a 13B parameter model is trained on over 20 times more training data than a 600M model, both can be backdoored by the same small number of poisoned documents. Our results challenge the common assumption that attackers need to control a percentage of train
Explore this link on the map →saved by
related reading
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samplesarxiv.org
- Phantom Transfer and the Basic Science of Data Poisoning — LessWronglesswrong.com
- [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoningarxiv.org
- Nicholas Carlininicholas.carlini.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Large Language Diffusion Modelsarxiv.org
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- HackAPromptpaper.hackaprompt.com
- Advice for making robust-to-training model organismsblog.redwoodresearch.org
- [2505.04741] When Bad Data Leads to Good Modelsarxiv.org