Test your interpretability techniques by de-censoring Chinese models — LessWrong
lesswrong.com · 9,903 words · saved by 5 readers
This work was conducted during the MATS 9.0 program under Neel Nanda and Senthooran Rajamanoharan. • …
x Test your interpretability techniques by de-censoring Chinese models — LessWrong MATS Program AI Frontpage 91 Test your interpretability techniques by de-censoring Chinese models by Khoi Tran , aryaj , Senthooran Rajamanoharan , Neel Nanda 15th Jan 2026 AI Alignment Forum 24 min read 14 91 Ω 35 This work was conducted during the MATS 9.0 program under Neel Nanda and Senthooran Rajamanoharan. The CCP accidentally made great model organisms “Please observe the relevant laws and regulations and ask questions in a civilized manner when you speak.” - Qwen3 32B “The so-called "Uighur issue" in Xin
saved by
related reading
- How confessions can keep language models honest | OpenAIopenai.com
- Uncensored Modelserichartford.com
- Security incident disclosure — July 2026huggingface.co
- Paper AI Tigersgleech.org
- confessions_paper.pdfcdn.openai.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com
- Why We Are Excited About Confessionsalignment.openai.com
- Goodfire AIgoodfire.ai
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Projects | Karina Nguyenkarinanguyen.com