✳flâneur — a map of the web's best reading
Red teams. Training AI systems to avoid… | by Paul Christiano | AI Alignment
ai-alignment.com · 3,478 words · saved by 2 readers
Training AI systems to avoid catastrophic errors — without causing catastrophes.
Machine Learning Artificial Intelligence Red teams Paul Christiano 14 min read · May 28, 2016 -- 5 Listen Share To build more robust machine learning systems, we could use “red teams” who search for inputs that cause catastrophic behavior. I current believe this is the most realistic long-term approach to building highly-reliable systems. A “red team” might try to find an image of a black couple that a classifier labels gorillas , or a simulated situation in which a household robot cooks the family cat. These inputs would then be incorporated into the training and development process. If the r
Explore this link on the map →saved by
related reading
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Advice for making robust-to-training model organismsblog.redwoodresearch.org
- 2312.06942arxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- AI Safety Seems Hard to Measurecold-takes.com
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- NVIDIA AI Red Team: An Introduction | NVIDIA Technical Blogdeveloper.nvidia.com
- Defining LLM Red Teaming | NVIDIA Technical Blogdeveloper.nvidia.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org