Race and Gender Bias As An Example of Unfaithful Chain of Thought in the Wild — LessWrong
Summary: We found that LLMs exhibit significant race and gender bias in realistic hiring scenarios, but their chain-of-thought reasoning shows zero evidence of this bias. This serves as a nice example of a 100% unfaithful CoT "in the wild" where the LLM strongly suppresses the unfaithful behavior. We also find that interpretability-based interventions succeeded while prompting failed, suggesting this may be an example of interpretability being the best practical tool for a real world problem. For context on our paper, the tweet thread is here and the paper is here. Chain of Thought (CoT) monitoring has emerged as a popular research area in AI safety. The idea is simple - have the AIs reason in English text when solving a problem, and monitor the reasoning for misaligned behavior. For example, OpenAI recently published a paper on using CoT monitoring to detect reward hacking during RL. An obvious concern is that the CoT may not be faithful to the model’s reasoning. Several papers have
x Race and Gender Bias As An Example of Unfaithful Chain of Thought in the Wild — LessWrong Chain-of-Thought Alignment Interpretability (ML & AI) Language Models (LLMs) AI Frontpage 2025 Top Fifty: 14 % 191 Race and Gender Bias As An Example of Unfaithful Chain of Thought in the Wild by Adam Karvonen , Sam Marks 2nd Jul 2025 4 min read 26 191 Summary: We found that LLMs exhibit significant race and gender bias in realistic hiring scenarios, but their chain-of-thought reasoning shows zero evidence of this bias. This serves as a nice example of a 100% unfaithful CoT "in the wild" where the LLM s
Explore this link on the map →related reading
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- the case for CoT unfaithfulness is overstated — LessWronglesswrong.com
- [2305.04388] Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Promptingarxiv.org
- The Unintelligibility is Ours: Notes on Chain of Thought1a3orn.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- [2405.18915] Towards Faithful Chain-of-Thought: Large Language Models are Bridging Reasonersar5iv.labs.arxiv.org
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Bias and limitations · Hugging Facehuggingface.co
- [2407.12856] AI-AI Bias: large language models favor communications generated by large language modelsarxiv.org
- Towards Faithful Chain-of-Thought: Large Language Models are Bridging Reasonersarxiv.org
- Measuring Faithfulness in Chain-of-Thought Reasoning \ Anthropicanthropic.com