Red teams. Training AI systems to avoid… | by Paul Christiano | AI Alignment
ai-alignment.com · 3,478 words · saved by 2 readers
Training AI systems to avoid catastrophic errors — without causing catastrophes.
Machine Learning Artificial Intelligence Red teams Paul Christiano 14 min read · May 28, 2016 -- 5 Listen Share To build more robust machine learning systems, we could use “red teams” who search for inputs that cause catastrophic behavior. I current believe this is the most realistic long-term approach to building highly-reliable systems. A “red team” might try to find an image of a black couple that a classifier labels gorillas , or a simulated situation in which a household robot cooks the family cat. These inputs would then be incorporated into the training and development process. If the r
saved by
related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Advice for making robust-to-training model organismsblog.redwoodresearch.org
- GitHub - requie/AI-Red-Teaming-Guide: A comprehensive guide to adversarial testing and security evaluation of AI systems, helping organizations identify vulnerabilities before attackers exploit them.github.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- 2312.06942arxiv.org
- Recommendations-for-Using-Red-Teaming-for-AI-Accountability-PolicyBrief.pdfdatasociety.net
- How our new Control Red Team is stress-testing frontier monitors | AISI Workaisi.gov.uk
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- What failure looks like — AI Alignment Forumalignmentforum.org
- AI Safety Seems Hard to Measurecold-takes.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org