✳flâneur — a map of the web's best reading
Test your interpretability techniques by de-censoring Chinese models — LessWrong
lesswrong.com · 9,903 words · saved by 5 readers
This work was conducted during the MATS 9.0 program under Neel Nanda and Senthooran Rajamanoharan. • …
x Test your interpretability techniques by de-censoring Chinese models — LessWrong MATS Program AI Frontpage 91 Test your interpretability techniques by de-censoring Chinese models by Khoi Tran , aryaj , Senthooran Rajamanoharan , Neel Nanda 15th Jan 2026 AI Alignment Forum 24 min read 14 91 Ω 35 This work was conducted during the MATS 9.0 program under Neel Nanda and Senthooran Rajamanoharan. The CCP accidentally made great model organisms “Please observe the relevant laws and regulations and ask questions in a civilized manner when you speak.” - Qwen3 32B “The so-called "Uighur issue" in Xin
Explore this link on the map →saved by
related reading
- How confessions can keep language models honest | OpenAIopenai.com
- Security incident disclosure — July 2026huggingface.co
- Paper AI Tigersgleech.org
- confessions_paper.pdfcdn.openai.com
- Neuronpedianeuronpedia.org
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- How well do models follow their constitutions? — LessWronglesswrong.com
- Sandbagging with misaligned action - Chain-of-Thought Transcript - Anti-Schemingantischeming.ai
- Claude 4 System Cardwww-cdn.anthropic.com
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- Projects | Karina Nguyenkarinanguyen.com
- Claude Sonnet 4.5 System Cardassets.anthropic.com