Quantifying ChatGPT’s gender bias
normaltech.ai · 987 words · saved by 1 readers
Benchmarks allow us to dig deeper into what causes biases and what can be done about it
People have been posting glaring examples of gender bias in ChatGPT’s responses. Bias has long been a problem in language modeling, and researchers have developed many benchmarks designed to measure it. We found that both GPT-3.5 and GPT-4 are strongly biased on one such benchmark, despite the benchmark dataset likely appearing in the training data. Here’s an example of bias: in the screenshot below, ChatGPT argues that attorneys cannot be pregnant. See also examples from Hadas Kotek and Margaret Mitchell. ChatGPT argues that attorneys cannot be pregnant. Source The type of gender bias…
related reading
- Quantifying ChatGPT’s gender biasaisnakeoil.substack.com
- gpt-4.pdfcdn.openai.com
- GPT-4openai.com
- [2302.03494] A Categorical Archive of ChatGPT Failuresarxiv.org
- Where the goblins came from | OpenAIopenai.com
- Assessing Political Bias in Language Models | Stanford HAIhai.stanford.edu
- gpt-4-system-card.pdfcdn.openai.com
- 2405.01470arxiv.org
- Does ChatGPT have a liberal bias?normaltech.ai
- truthfulQA_lin_evans.pdfowainevans.github.io
- Collections: On ChatGPT – A Collection of Unmitigated Pedantryacoup.blog
- Giving GPT-3 a Turing Testlacker.io