2408.12798
arxiv.org · 7,136 words · saved by 1 readers
N/A
BACKDOOR LLM: A Comprehensive Benchmark for Backdoor Attacks and Defenses on Large Language Models Yige Li1 , Hanxun Huang2 , Yunhan Zhao3 , Xingjun Ma3 , Jun Sun1 arXiv:2408.12798v2 [cs.AI] 19 May 2025 1 Singapore Management University 2 The University of Melbourne 3 Fudan University…
saved by
related reading
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samplesarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Nicholas Carlininicholas.carlini.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired Contentarxiv.org
- Productizing Large Language Modelsblog.replit.com
- Large Language Diffusion Modelsarxiv.org
- [2602.04899] Phantom Transfer: Data-level Defences are Insufficient Against Data Poisoningarxiv.org