flâneur — a map of the web's best reading

Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases

alignment.anthropic.com · 3,966 words · saved by 4 readers

TL;DR: We provide some evidence that Claude 3.7 Sonnet doesn’t encode hidden reasoning in its scratchpad by showing that training it to use paraphrased versions of the scratchpads does not degrade performance. The scratchpads from reasoning models look human understandable: when reasoning about a math problem, reasoning models consider intermediate steps similar to the ones I would use, backtrack and double-check their work as I would. But even if scratchpads look like they perform human-like reasoning, scratchpads might improve performance through some less human-understandable mechanisms. One particularly worrying possibility is that models could encode additional reasoning in syntax of the text (e.g. encoding a bit in using a bulleted vs a numbered list, and then using this bit later in the scratchpad). This is sometimes called encoded reasoning or Chain-of-Thought steganography. If LLMs learned how to use encoded reasoning during RL, they might be able to use it in deployment to re

Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases Alignment Science Blog Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases TL;DR: We provide some evidence that Claude 3.7 Sonnet doesn’t encode hidden reasoning in its scratchpad by showing that training it to use paraphrased versions of the scratchpads does not degrade performance. The scratchpads from reasoning models look human understandable: when reasoning about a math problem, reasoning models consider intermediate steps similar to the ones I would use, backtra

Explore this link on the map →

saved by

related reading