✳flâneur — a map of the web's best reading
Discovering Backdoor Triggers — LessWrong
lesswrong.com · 5,597 words · saved by 1 readers
Authors: Andrew Qin*, Tim Hua*, Samuel Marks, Arthur Conmy, Neel Nanda …
x Discovering Backdoor Triggers — LessWrong Interpretability (ML & AI) MATS Program AI Frontpage 57 Discovering Backdoor Triggers by andrq , Tim Hua , Sam Marks , Arthur Conmy , Neel Nanda 19th Aug 2025 AI Alignment Forum 16 min read 4 57 Ω 24 Authors: Andrew Qin*, Tim Hua*, Samuel Marks, Arthur Conmy, Neel Nanda Andrew and Tim are co-first authors. This is a research progress report from Neel Nanda’s MATS 8.0 stream. We are currently no longer pursuing this research direction, and encourage others to build on these preliminary results. tl;dr. We study whether we can reverse-engineer the trigg
Explore this link on the map →saved by
related reading
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- 2312.06942arxiv.org
- Backdoors have universal representations across large language models — LessWronglesswrong.com
- secret-loyalties-whitepaper.pdfformationresearch.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Neuronpedianeuronpedia.org
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Auditing language models for hidden objectives — LessWronglesswrong.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org