GPT-Red: Unlocking Self-Improvement for Robustness | OpenAI
Red-teaming is essential to discovering vulnerabilities and improving the robustness of our models. However, current approaches are not scalable, creating a bottleneck. Commonly used robustness evaluations have already been saturated by our latest models. We need to develop methods that allow safety and alignment to scale alongside model capabilities. What we did We trained GPT‑Red, an automated red-teaming model that scales our ability to find vulnerabilities so we can fix them before wider deployment. GPT‑Red is a strong red-teamer, and our previous models are highly vulnerable to its prompt injection attacks. We use GPT‑Red to adversarially train GPT‑5.6, making it much more robust to prompt injections. We will continue to scale this approach alongside human and third-party red-teaming, layered safeguards, and real-time monitoring. AI systems commonly encounter third-party data through browsers, connected apps, local files, and other tools. These affordances are necessary for perfor