Assessing Political Bias in Language Models
hai.stanford.edu · 1,109 words · saved by 1 readers
Researchers develop a new tool to measure how well popular large language models align with public opinion to evaluate bias in chatbots.
DALL-E Researchers develop a new tool to measure how well popular large language models align with public opinion to evaluate bias in chatbots. The language models behind ChatGPT and other generative AI are trained on written words that have been culled from libraries, scraped from websites and social media, and pulled from news reports and speech transcripts from across the world. There are 250 billion such words behind GPT-3.5, the model fueling ChatGPT, for instance, and GPT-4 is now here. Now new research from Stanford University has quantified exactly how well (or, actually, how poorly) t
related reading
- [2303.17548] Whose Opinions Do Language Models Reflect?arxiv.org
- [2306.13000] Apolitical Intelligence? Auditing Delphi's responses on controversial political issues in the USarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- 2308.03958arxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Cognitive Biases in Large Language Models — LessWronglesswrong.com
- Projects | Karina Nguyenkarinanguyen.com
- Unsupervised Elicitationalignment.anthropic.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Training language models to follow instructions with human feedback.pdfproceedings.neurips.cc
- Does ChatGPT have a liberal bias?normaltech.ai