Reverse engineering Claude's CVE-2026-2796 exploit
Today we published an update on our collaboration with Mozilla, in which Claude Opus 4.6 found 22 vulnerabilities in Firefox over the course of two weeks. As part of that work, we evaluated whether Claude could go further: exploit the bugs, as well as find them. This blog post will deep dive into how Claude wrote an exploit for CVE-2026-2796 (now patched). This is another data point for the trajectory of LLM’s cyber capabilities. In September, we noted that Claude's success rate on Cybench had doubled in six months. In early February we demonstrated that Claude’s success rate on Cybergym doubled in four months. We’re sharing this case study to provide an early glimpse into what we expect will be LLMs’ improving ability to author exploits. To be clear, the exploit that Claude wrote only works within a testing environment that intentionally removes some of the security features of modern web browsers. Claude isn't yet writing “full-chain” exploits that combine multiple vulnerabilities t
Frontier Red Team Reverse engineering Claude's CVE-2026-2796 exploit Mar 6, 2026 Evyatar Ben Asher, Keane Lucas, Nicholas Carlini, Newton Cheng, and Daniel Freeman Introduction Today we published an update on our collaboration with Mozilla, in which Claude Opus 4.6 found 22 vulnerabilities in Firefox over the course of two weeks. As part of that work, we evaluated whether Claude could go further: exploit the bugs, as well as find them. This blog post will deep dive into how Claude wrote an exploit for CVE-2026-2796 (now patched). This is another data point for the trajectory of LLM’s cyber cap
Explore this link on the map →related reading
- Assessing Claude Mythos Preview’s cybersecurity capabilities \ Anthropicred.anthropic.com
- All learning materials - detailed | Web Security Academyportswigger.net
- LLM-discovered 0 days \ Anthropicred.anthropic.com
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- Project Glasswing: Securing critical software for the AI era \ Anthropicanthropic.com
- Measuring LLMs' impact on N-day exploits \ Anthropicred.anthropic.com
- Rewriting Bun in Rust | Bun Blogbun.com
- Best practices for Claude Code - Claude Code Docsanthropic.com
- Zenbleedlock.cmpxchg8b.com
- From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code - Project Zerogoogleprojectzero.blogspot.com
- Security incident disclosure — July 2026huggingface.co
- Finding Miscompiles for Fun, Not Profit - by Justin Lebarnewsletter.semianalysis.com