✳flâneur — a map of the web's best reading
Project Zero: From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code
googleprojectzero.blogspot.com · 2,484 words · saved by 1 readers
Posted by the Big Sleep team Introduction In our previous post, Project Naptime: Evaluating Offensive Security Capabilities of Large L...
Posted by the Big Sleep team Introduction In our previous post, Project Naptime: Evaluating Offensive Security Capabilities of Large Language Models , we introduced our framework for large-language-model-assisted vulnerability research and demonstrated its potential by improving the state-of-the-art performance on Meta's CyberSecEval2 benchmarks. Since then, Naptime has evolved into Big Sleep, a collaboration between Google Project Zero and Google DeepMind. Today, we're excited to share the first real-world vulnerability discovered by the Big Sleep agent : an exploitable stack buffer underflow
Explore this link on the map →related reading
- From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code - Project Zerogoogleprojectzero.blogspot.com
- Assessing Claude Mythos Preview’s cybersecurity capabilities \ Anthropicred.anthropic.com
- Project Glasswing: Securing critical software for the AI era \ Anthropicanthropic.com
- LLM-discovered 0 days \ Anthropicred.anthropic.com
- Mythos finds a curl vulnerability | daniel.haxx.sedaniel.haxx.se
- Measuring LLMs’ ability to develop exploits \ Anthropicred.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Measuring LLMs' impact on N-day exploits \ Anthropicred.anthropic.com
- Finding Miscompiles for Fun, Not Profit - by Justin Lebarnewsletter.semianalysis.com
- [2312.12575] LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarksarxiv.org
- Zenbleedlock.cmpxchg8b.com
- Simon Willison’s Weblogsimonwillison.net