Training Qwen-1.5B with a CoT legibility penalty — LessWrong
I tried training Qwen2.5-1.5B with RL on math to both get correct answers and have a CoT that doesn’t look like human-understandable math reasoning. RL sometimes succeeds at hacking my monitor, and when I strengthen my monitor, it fails at finding CoT that are both illegible and helpful, even after training for roughly 4000 steps (~1B generated tokens). Exploring into obfuscated reasoning is hard! These results were also released in the Appendix of Training fails to elicit subtle reasoning in current language models. Chain-of-Thoughts (CoTs) can help reason for many more serial steps than there are layers in a Transformer. But one worry is that LLMs might hide their real reasoning in a plausible benign CoT. Previous work has demonstrated that in toy setups, LLMs can learn extremely simple encodings, but nothing sufficiently general to e.g. be helpful to solve a wide range of math problems. To find naturally emerging encodings, I relax the “plausible benign” constraint and try to find m
x Training Qwen-1.5B with a CoT legibility penalty — LessWrong AI Frontpage 68 Training Qwen-1.5B with a CoT legibility penalty by Fabien Roger 9th Oct 2025 5 min read 7 68 I tried training Qwen2.5-1.5B with RL on math to both get correct answers and have a CoT that doesn’t look like human-understandable math reasoning. RL sometimes succeeds at hacking my monitor, and when I strengthen my monitor, it fails at finding CoT that are both illegible and helpful, even after training for roughly 4000 steps (~1B generated tokens). Exploring into obfuscated reasoning is hard! These results were also re
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- DeepSeek-R1arxiv.org
- [2510.27338] Reasoning Models Sometimes Output Illegible Chains of Thoughtarxiv.org
- Learning to reason with LLMs | OpenAIopenai.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- [Research Note] Optimizing The Final Output Can Obfuscate CoT — LessWronglesswrong.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com
- Vestigial reasoning in RL — LessWronglesswrong.com
- Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrasesalignment.anthropic.com
- Chain-of-Thought Promptinglearnprompting.org