flâneur — a map of the web's best reading

Race and Gender Bias As An Example of Unfaithful Chain of Thought in the Wild — LessWrong

lesswrong.com · 5,452 words · saved by 1 readers

Summary: We found that LLMs exhibit significant race and gender bias in realistic hiring scenarios, but their chain-of-thought reasoning shows zero evidence of this bias. This serves as a nice example of a 100% unfaithful CoT "in the wild" where the LLM strongly suppresses the unfaithful behavior. We also find that interpretability-based interventions succeeded while prompting failed, suggesting this may be an example of interpretability being the best practical tool for a real world problem. For context on our paper, the tweet thread is here and the paper is here. Chain of Thought (CoT) monitoring has emerged as a popular research area in AI safety. The idea is simple - have the AIs reason in English text when solving a problem, and monitor the reasoning for misaligned behavior. For example, OpenAI recently published a paper on using CoT monitoring to detect reward hacking during RL. An obvious concern is that the CoT may not be faithful to the model’s reasoning. Several papers have

x Race and Gender Bias As An Example of Unfaithful Chain of Thought in the Wild — LessWrong Chain-of-Thought Alignment Interpretability (ML & AI) Language Models (LLMs) AI Frontpage 2025 Top Fifty: 14 % 191 Race and Gender Bias As An Example of Unfaithful Chain of Thought in the Wild by Adam Karvonen , Sam Marks 2nd Jul 2025 4 min read 26 191 Summary: We found that LLMs exhibit significant race and gender bias in realistic hiring scenarios, but their chain-of-thought reasoning shows zero evidence of this bias. This serves as a nice example of a 100% unfaithful CoT "in the wild" where the LLM s

Explore this link on the map →

related reading