flâneur

Backdoors have universal representations across large language models — LessWrong

lesswrong.com · 5,326 words · saved by 1 readers

by Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Amirali Abdullah • …

x Backdoors have universal representations across large language models — LessWrong Deceptive Alignment Interpretability (ML & AI) AI Frontpage 18 Backdoors have universal representations across large language models by Amirali Abdullah , Narmeen , Dhruv Nathawani , nirmalendu prakash 6th Dec 2024 20 min read 0 18 by Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash, Amirali Abdullah This work was done by Narmeen Oozeer as a research fellow at Martian, under an AI safety grant supervised by PIs Amirali Abdullah and Dhruv Nathawani. Special thanks to Sasha Hydrie, Chaithanya Bandi and Shriyas

related reading