✳flâneur — a map of the web's best reading
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
arxiv.org · 15,310 words · saved by 1 readers
N/A
# link_17i16eksnws.pdf ## Metadata - PDFFormatVersion=1.7 - IsLinearized=false - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - Author=Alexandra Souly; Javier Rando; Ed Chapman; Xander Davies; Burak Hasircioglu; Ezzeldin Shereen; Carlos Mougan; Vasilios Mavroudis; Erik Jones; Chris Hicks; Nicholas Carlini; Yarin Gal; Robert Kirk - Creator=arXiv GenPDF (tex2pdf:e76afa9) - Custom.DOI=https://doi.org/10.48550/arXiv.2510.07192 - Custom.License=http://creativecommons.org/licenses/by/4.0/ - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.14159
Explore this link on the map →saved by
related reading
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoningarxiv.org
- Phantom Transfer and the Basic Science of Data Poisoning — LessWronglesswrong.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- Nicholas Carlininicholas.carlini.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- Large Language Diffusion Modelsarxiv.org
- Pathways Language Model (PaLM): Scaling to 540 Billion Parameters for Breakthrouai.googleblog.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io