flâneur — a map of the web's best reading

Do models say what they learn? — LessWrong

lesswrong.com · 9,460 words · saved by 1 readers

This is a writeup of preliminary research studying whether models verbalize what they learn during RL training. This research is incomplete, and not up to the rigorous standards of a publication. We're sharing our progress so far, and would be happy for others to further explore this direction. Code to reproduce the core experiments is available here. This study investigates whether language models articulate new behaviors learned during reinforcement learning (RL) training. Specifically, we train a 7-billion parameter chat model on a loan approval task, creating datasets with simple biases (e.g., "approve all Canadian applicants") and training the model via RL to adopt these biases. We find that models learn to make decisions based entirely on specific attributes (e.g. nationality, gender) while rarely articulating these attributes as factors in their reasoning. Chain-of-thought (CoT) monitoring is one of the most promising methods for AI oversight. CoT monitoring can be used to track

x Do models say what they learn? — LessWrong AI Frontpage 2025 Top Fifty: 14 % 127 Do models say what they learn? by Andy Arditi , marvinli , Joe Benton , Miles Turpin 22nd Mar 2025 15 min read 12 127 This is a writeup of preliminary research studying whether models verbalize what they learn during RL training. This research is incomplete, and not up to the rigorous standards of a publication. We're sharing our progress so far, and would be happy for others to further explore this direction. Code to reproduce the core experiments is available here . Summary This study investigates whether lang

Explore this link on the map →

related reading