[2602.09416] Are Language Models Sensitive to Morally Irrelevant Distractors?
Abstract:With the rapid uptake of large language models (LLMs) across high-stakes settings, it is becoming increasingly important to ensure that LLMs behave in ways that align with human values. Existing moral benchmarks for this purpose often prompt LLMs with value statements, moral scenarios, or psychological questionnaires, with the implicit underlying assumption that LLMs report somewhat stable moral preferences. However, moral psychology research has shown that even human moral judgements are sensitive to morally irrelevant situational factors such as the smell of cinnamon rolls or the level of ambient noise, thereby challenging moral theories which assume that human moral judgements are stable. Here we draw inspiration from this "situationist" view of moral psychology to evaluate whether LLMs exhibit similar cognitive moral biases. We curate a novel multimodal dataset of 60 "moral distractors" from existing psychological datasets of emotionally-valenced images and narratives, which have no moral relevance to the situation presented. After injecting these distractors into existing moral benchmarks, we find that moral distractors can shift the moral judgements of LLMs by over 30% even in unambiguous scenarios, highlighting the instability of LLMs' moral judgements and the need for more contextual approaches to AI alignment.
[2602.09416] Are Language Models Sensitive to Morally Irrelevant Distractors? Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2602.09416 (cs) [Submitted on 10 Feb 2026 ( v1 ), last revised 20 Jun 2026 (this version, v2)] Title: Are Language Models Sensitive to Morally Irrelevant Distractors? Authors: Andrew Shaw , Christina Hahn , Catherine Rasgaitis , Yash Mishra , Alisa Liu , Natasha Jaques , Yulia Tsvetkov , Amy X. Zhang View a PDF of the paper tit
related reading
- [2405.17345] Exploring and steering the moral compass of Large Language Modelsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- [2303.17548] Whose Opinions Do Language Models Reflect?arxiv.org
- LMCA_dataset.pdfandrew.cmu.edu
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesarxiv.org
- Cognitive Biases in Large Language Models — LessWronglesswrong.com
- 2308.03958arxiv.org
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesarxiv.org
- Language Models Learn to Mislead Humans via RLHFarxiv.org
- [2503.03750] The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systemsarxiv.org