Quantifying ChatGPT’s gender bias
aisnakeoil.substack.com · 971 words · saved by 1 readers
Benchmarks allow us to dig deeper into what causes biases and what can be done about it
People have been posting glaring examples of gender bias in ChatGPT’s responses. Bias has long been a problem in language modeling, and researchers have developed many benchmarks designed to measure it. We found that both GPT-3.5 and GPT-4 are strongly biased on one such benchmark, despite the benchmark dataset likely appearing in the training data. Here’s an example of bias: in the screenshot below, ChatGPT argues that attorneys cannot be pregnant. See also examples from Hadas Kotek and Margaret Mitchell. ChatGPT argues that attorneys cannot be pregnant. Source The type of gender bias…
related reading
- Quantifying ChatGPT’s gender biasnormaltech.ai
- gpt-4.pdfcdn.openai.com
- GPT-4openai.com
- Where the goblins came from | OpenAIopenai.com
- [2302.03494] A Categorical Archive of ChatGPT Failuresarxiv.org
- gpt-4-system-card.pdfcdn.openai.com
- Assessing Political Bias in Language Models | Stanford HAIhai.stanford.edu
- 2405.01470arxiv.org
- Does ChatGPT have a liberal bias?normaltech.ai
- truthfulQA_lin_evans.pdfowainevans.github.io
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Giving GPT-3 a Turing Testlacker.io