[2310.18512] Preventing Language Models From Hiding Their Reasoning
Large language models (LLMs) often benefit from intermediate steps of reasoning to generate answers to complex problems. When these intermediate steps of reasoning are used to monitor the activity of the model, it is e…
Preventing Language Models From Hiding Their Reasoning Fabien Roger ∗ Ryan Greenblatt Redwood Research Abstract Large language models (LLMs) often benefit from intermediate steps of reasoning to generate answers to complex problems. When these intermediate steps of reasoning are used to monitor the activity of the model, it is essential that this explicit reasoning is faithful, i.e. that it reflects what the model is actually reasoning about. In this work, we focus on one potential way intermediate steps of reasoning could be unfaithful: encoded reasoning, where an LLM could encode intermediat
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Learning to reason with LLMs | OpenAIopenai.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrasesalignment.anthropic.com
- 2025: The year in LLMssimonwillison.net
- By Default, GPTs Think In Plain Sight — LessWronglesswrong.com
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- [2510.24941] Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thoughtarxiv.org
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- Taking LLMs Seriously (As Language Models) — LessWronglesswrong.com