flâneur

Quantifying ChatGPT’s gender bias

aisnakeoil.substack.com · 971 words · saved by 1 readers

Benchmarks allow us to dig deeper into what causes biases and what can be done about it

People have been posting glaring examples of gender bias in ChatGPT’s responses. Bias has long been a problem in language modeling, and researchers have developed many benchmarks designed to measure it. We found that both GPT-3.5 and GPT-4 are strongly biased on one such benchmark, despite the benchmark dataset likely appearing in the training data. Here’s an example of bias: in the screenshot below, ChatGPT argues that attorneys cannot be pregnant. See also examples from Hadas Kotek and Margaret Mitchell. ChatGPT argues that attorneys cannot be pregnant. Source The type of gender bias…

related reading