Project Zero: From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code
googleprojectzero.blogspot.com · 2,484 words · saved by 1 readers
Posted by the Big Sleep team Introduction In our previous post, Project Naptime: Evaluating Offensive Security Capabilities of Large L...
Posted by the Big Sleep team Introduction In our previous post, Project Naptime: Evaluating Offensive Security Capabilities of Large Language Models , we introduced our framework for large-language-model-assisted vulnerability research and demonstrated its potential by improving the state-of-the-art performance on Meta's CyberSecEval2 benchmarks. Since then, Naptime has evolved into Big Sleep, a collaboration between Google Project Zero and Google DeepMind. Today, we're excited to share the first real-world vulnerability discovered by the Big Sleep agent : an exploitable stack buffer underflow
related reading
- From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code - Project Zerogoogleprojectzero.blogspot.com
- Assessing Claude Mythos Preview’s cybersecurity capabilities \ Anthropicred.anthropic.com
- Hōrōshi バガボンド (@KatanaLarp) on Xx.com
- Project Glasswing: Securing critical software for the AI era \ Anthropicanthropic.com
- How we tracked down a 16-year-old SQLite bugtailscale.com
- LLM-discovered 0 days \ Anthropicred.anthropic.com
- Mythos finds a curl vulnerability | daniel.haxx.sedaniel.haxx.se
- Measuring LLMs’ ability to develop exploits \ Anthropicred.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Measuring LLMs' impact on N-day exploits \ Anthropicred.anthropic.com
- Using LLM-based Verification to Eliminate Bugs in Linux's Network Stackbasis.ai
- Finding Miscompiles for Fun, Not Profit - by Justin Lebarnewsletter.semianalysis.com