Do models say what they learn? — LessWrong
This is a writeup of preliminary research studying whether models verbalize what they learn during RL training. This research is incomplete, and not up to the rigorous standards of a publication. We're sharing our progress so far, and would be happy for others to further explore this direction. Code to reproduce the core experiments is available here. This study investigates whether language models articulate new behaviors learned during reinforcement learning (RL) training. Specifically, we train a 7-billion parameter chat model on a loan approval task, creating datasets with simple biases (e.g., "approve all Canadian applicants") and training the model via RL to adopt these biases. We find that models learn to make decisions based entirely on specific attributes (e.g. nationality, gender) while rarely articulating these attributes as factors in their reasoning. Chain-of-thought (CoT) monitoring is one of the most promising methods for AI oversight. CoT monitoring can be used to track
x Do models say what they learn? — LessWrong AI Frontpage 2025 Top Fifty: 14 % 127 Do models say what they learn? by Andy Arditi , marvinli , Joe Benton , Miles Turpin 22nd Mar 2025 15 min read 12 127 This is a writeup of preliminary research studying whether models verbalize what they learn during RL training. This research is incomplete, and not up to the rigorous standards of a publication. We're sharing our progress so far, and would be happy for others to further explore this direction. Code to reproduce the core experiments is available here . Summary This study investigates whether lang
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- DeepSeek-R1arxiv.org
- Vestigial reasoning in RL — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- o1 and Reasoning | AndoLogsblog.ando.ai
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- [2510.27338] Reasoning Models Sometimes Output Illegible Chains of Thoughtarxiv.org