[2604.21751] Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
Abstract:LLMs have been showing limitations when it comes to cultural coverage and competence, and in some cases show regional biases such as amplifying Western and Anglocentric viewpoints. While there have been works analysing the cultural capabilities of LLMs, there has not been specific work on highlighting LLM regional preferences when it comes to cultural-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ). The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan. Moveover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs and show less inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training.
Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs * Joseba Fernandez de Landa1 Carla Perez-Almendros2 Jose Camacho-Collados2 1 HiTZ Center - Ixa, University of the Basque Country EHU 2…
saved by
related reading
- What language do language models speak?tcz.hu
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- [2303.17548] Whose Opinions Do Language Models Reflect?arxiv.org
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Models are culturesrobknight.org.uk
- Thoughts on AI in academiatheinfinitesimal.substack.com
- 2025: The year in LLMssimonwillison.net
- Writing for LLMs So They Listen · Gwern.netgwern.net
- The two types of LLM preferencesnewsletter.danielpaleka.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Things we learned about LLMs in 2024simonwillison.net
- Creativity Has Left the Chat: The Price of Debiasing Language Modelsarxiv.org