0-Days \ red.anthropic.com
Nicholas Carlini*, Keane Lucas*, Evyatar Ben Asher*, Newton Cheng, Hasnain Lakhani, and David Forsythe *indicates equal contribution Claude Opus 4.6, released today, continues a trajectory of meaningful improvements in AI models’ cybersecurity capabilities. Last fall, we wrote that we believed we were at an inflection point for AI's impact on cybersecurity—that progress could become quite fast, and now was the moment to accelerate defensive use of AI. The evidence since then has only reinforced that view. AI models can now find high-severity vulnerabilities at scale. Our view is this is a moment to move quickly—to empower defenders and secure as much code as possible while the window exists. Opus 4.6 is notably better at finding high-severity vulnerabilities than previous models and a sign of how quickly things are moving. Security teams have been automating vulnerability discovery for years, investing heavily in fuzzing infrastructure and custom harnesses to find bugs at scale. But wh
Frontier Red Team Evaluating and mitigating the growing risk of LLM-discovered 0-days Feb 5, 2026 Nicholas Carlini * , Keane Lucas * , Evyatar Ben Asher * , Newton Cheng, Hasnain Lakhani, David Forsythe, and Kyla Guru * indicates equal contribution Claude Opus 4.6, released today , continues a trajectory of meaningful improvements in AI models’ cybersecurity capabilities. Last fall, we wrote that we believed we were at an inflection point for AI's impact on cybersecurity —that progress could become quite fast, and now was the moment to accelerate defensive use of AI. The evidence since then ha
Explore this link on the map →related reading
- Project Glasswing: Securing critical software for the AI era \ Anthropicanthropic.com
- Assessing Claude Mythos Preview’s cybersecurity capabilities \ Anthropicred.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Vulnerability Research Is Cooked - Quarrelsomesockpuppet.org
- Mythos finds a curl vulnerability | daniel.haxx.sedaniel.haxx.se
- Measuring LLMs' impact on N-day exploits \ Anthropicred.anthropic.com
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- Claude 4 System Cardwww-cdn.anthropic.com
- From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code - Project Zerogoogleprojectzero.blogspot.com
- [2312.12575] LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarksarxiv.org
- Nicholas Carlininicholas.carlini.com
- FIRST Mid-Year Vulnerability Forecast Confirms Historic Surge, Projects ~66,000 CVEs in 2026first.org