the case for CoT unfaithfulness is overstated — LessWrong
[Quickly written, unpolished. Also, it's possible that there's some more convincing work on this topic that I'm unaware of – if so, let me know. Also…
x the case for CoT unfaithfulness is overstated — LessWrong Chain-of-Thought Alignment AI Rationality Frontpage 271 the case for CoT unfaithfulness is overstated by nostalgebraist 29th Sep 2024 AI Alignment Forum 14 min read 45 271 Ω 110 [Quickly written, unpolished. Also, it's possible that there's some more convincing work on this topic that I'm unaware of – if so, let me know. Also also, it's possible I'm arguing with an imaginary position here and everyone already agrees with everything below.] In research discussions about LLMs, I often pick up a vibe of casual, generalized skepticism abo
Explore this link on the map →saved by
related reading
- [2405.18915] Towards Faithful Chain-of-Thought: Large Language Models are Bridging Reasonersar5iv.labs.arxiv.org
- Measuring Faithfulness in Chain-of-Thought Reasoning \ Anthropicanthropic.com
- Towards Faithful Chain-of-Thought: Large Language Models are Bridging Reasonersarxiv.org
- [2305.04388] Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Promptingarxiv.org
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- [2505.05410] Reasoning Models Don't Always Say What They Thinkarxiv.org
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- What’s your AI thinking? - AI Digesttheaidigest.org
- How confessions can keep language models honest | OpenAIopenai.com
- The fragile foundations of CoT monitoring | Christopher Pottsweb.stanford.edu