✳flâneur — a map of the web's best reading
Backdoors have universal representations across large language models — LessWrong
lesswrong.com · 5,326 words · saved by 1 readers
by Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Amirali Abdullah • …
x Backdoors have universal representations across large language models — LessWrong Deceptive Alignment Interpretability (ML & AI) AI Frontpage 18 Backdoors have universal representations across large language models by Amirali Abdullah , Narmeen , Dhruv Nathawani , nirmalendu prakash 6th Dec 2024 20 min read 0 18 by Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Amirali Abdullah This work was done by Narmeen Oozeer as a research fellow at Martian, under an AI safety grant supervised by PIs Amirali Abdullah and Dhruv Nathawani. Special thanks to Sasha Hydrie, Chaithanya Bandi and Shriyas
Explore this link on the map →related reading
- Transformer Circuits Threadtransformer-circuits.pub
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- Towards Monosemanticity: Decomposing Language Models With Dictionary Learningtransformer-circuits.pub
- Discovering Backdoor Triggers — LessWronglesswrong.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Mapping the Mind of a Large Language Model \ Anthropicanthropic.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samplesarxiv.org
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- I Trained a Language Model. Then I Built a Brain Scanner and Looked Inside It. | by Caleb DeLeeuw | Mediummedium.com
- Sparse Autoencoders Reveal Universal Feature Spaces Across Large Language Modelsarxiv.org