✳flâneur — a map of the web's best reading
Vedaangh Rungta
vedaangh.com · 1,410 words · saved by 1 readers
Personal Website
← Back Introduction Chinese open-weight and open-source–style LLMs are now easy to download, fine-tune, and self-host, and they dominate the open-source LLM landscape, where Western alternatives are few and far between. There has been little public work investigating the political alignment and institutional allegiances of these models, and the risks that may arise from deploying them in Western contexts. To make those risks more legible, we evaluated six Chinese frontier models across six behavioral evaluations. Models were evaluated on: (1) answering questions on matters sensitive to the CCP
Explore this link on the map →related reading
- Paper AI Tigersgleech.org
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Test your interpretability techniques by de-censoring Chinese models — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- China AI Bulletin 1 - by Emmie Hine - China AI Bulletinchinaaibulletin.substack.com
- 2025: The year in LLMssimonwillison.net
- Evaluating DeepSeek v4 Pro for Frontier Risks · Neo Researchneoresearch.ai
- How far does alignment midtraining generalize?alignment.openai.com
- Alignment faking in large language modelsarxiv.org
- State of AI 2025: 100T Token LLM Usage Study | OpenRouteropenrouter.ai
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com