flâneur — a map of the web's best reading

Training Qwen-1.5B with a CoT legibility penalty — LessWrong

lesswrong.com · 3,917 words · saved by 1 readers

I tried training Qwen2.5-1.5B with RL on math to both get correct answers and have a CoT that doesn’t look like human-understandable math reasoning. RL sometimes succeeds at hacking my monitor, and when I strengthen my monitor, it fails at finding CoT that are both illegible and helpful, even after training for roughly 4000 steps (~1B generated tokens). Exploring into obfuscated reasoning is hard! These results were also released in the Appendix of Training fails to elicit subtle reasoning in current language models. Chain-of-Thoughts (CoTs) can help reason for many more serial steps than there are layers in a Transformer. But one worry is that LLMs might hide their real reasoning in a plausible benign CoT. Previous work has demonstrated that in toy setups, LLMs can learn extremely simple encodings, but nothing sufficiently general to e.g. be helpful to solve a wide range of math problems. To find naturally emerging encodings, I relax the “plausible benign” constraint and try to find m

x Training Qwen-1.5B with a CoT legibility penalty — LessWrong AI Frontpage 68 Training Qwen-1.5B with a CoT legibility penalty by Fabien Roger 9th Oct 2025 5 min read 7 68 I tried training Qwen2.5-1.5B with RL on math to both get correct answers and have a CoT that doesn’t look like human-understandable math reasoning. RL sometimes succeeds at hacking my monitor, and when I strengthen my monitor, it fails at finding CoT that are both illegible and helpful, even after training for roughly 4000 steps (~1B generated tokens). Exploring into obfuscated reasoning is hard! These results were also re

Explore this link on the map →

related reading