flâneur — a map of the web's best reading

Backdoors have universal representations across large language models — LessWrong

lesswrong.com · 5,326 words · saved by 1 readers

by Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Amirali Abdullah • …

x Backdoors have universal representations across large language models — LessWrong Deceptive Alignment Interpretability (ML & AI) AI Frontpage 18 Backdoors have universal representations across large language models by Amirali Abdullah , Narmeen , Dhruv Nathawani , nirmalendu prakash 6th Dec 2024 20 min read 0 18 by Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Amirali Abdullah This work was done by Narmeen Oozeer as a research fellow at Martian, under an AI safety grant supervised by PIs Amirali Abdullah and Dhruv Nathawani. Special thanks to Sasha Hydrie, Chaithanya Bandi and Shriyas

Explore this link on the map →

related reading