Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases
TL;DR: We provide some evidence that Claude 3.7 Sonnet doesn’t encode hidden reasoning in its scratchpad by showing that training it to use paraphrased versions of the scratchpads does not degrade performance. The scratchpads from reasoning models look human understandable: when reasoning about a math problem, reasoning models consider intermediate steps similar to the ones I would use, backtrack and double-check their work as I would. But even if scratchpads look like they perform human-like reasoning, scratchpads might improve performance through some less human-understandable mechanisms. One particularly worrying possibility is that models could encode additional reasoning in syntax of the text (e.g. encoding a bit in using a bulleted vs a numbered list, and then using this bit later in the scratchpad). This is sometimes called encoded reasoning or Chain-of-Thought steganography. If LLMs learned how to use encoded reasoning during RL, they might be able to use it in deployment to re
Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases Alignment Science Blog Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases TL;DR: We provide some evidence that Claude 3.7 Sonnet doesn’t encode hidden reasoning in its scratchpad by showing that training it to use paraphrased versions of the scratchpads does not degrade performance. The scratchpads from reasoning models look human understandable: when reasoning about a math problem, reasoning models consider intermediate steps similar to the ones I would use, backtra
Explore this link on the map →saved by
related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- On the Biology of a Large Language Modeltransformer-circuits.pub
- Learning to reason with LLMs | OpenAIopenai.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- o1 and Reasoning | AndoLogsblog.ando.ai
- The Universe of Discourseblog.plover.com
- Reasoning as Trajectoriesslhleosun.github.io
- Reasoning models | OpenAI APIplatform.openai.com
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Claude's extended thinking \ Anthropicanthropic.com