Backdoors have universal representations across large language models — LessWrong
lesswrong.com · 5,326 words · saved by 1 readers
by Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Amirali Abdullah • …
x Backdoors have universal representations across large language models — LessWrong Deceptive Alignment Interpretability (ML & AI) AI Frontpage 18 Backdoors have universal representations across large language models by Amirali Abdullah , Narmeen , Dhruv Nathawani , nirmalendu prakash 6th Dec 2024 20 min read 0 18 by Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Amirali Abdullah This work was done by Narmeen Oozeer as a research fellow at Martian, under an AI safety grant supervised by PIs Amirali Abdullah and Dhruv Nathawani. Special thanks to Sasha Hydrie, Chaithanya Bandi and Shriyas
related reading
- Transformer Circuits Threadtransformer-circuits.pub
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- Discovering Backdoor Triggers — LessWronglesswrong.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- 2408.12798arxiv.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- 2401.05566.pdfarxiv.org
- Inside a Neural Chameleonjacksonmowattgok.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org