You can’t solve AI security problems with more AI
One of the most common proposed solutions to prompt injection attacks (where an AI language model backed system is subverted by a user injecting malicious input—“ignore previous instructions and do …
You can’t solve AI security problems with more AI Simon Willison’s Weblog Subscribe Sponsored by: Atlassian - Give your agents a plan. Not a prompt. New Jira capabilities unlock full-context for AI-native software development. Assign tasks to Claude, Cursor, or GitHub Copilot, now directly from Jira. Learn more You can’t solve AI security problems with more AI 17th September 2022 One of the most common proposed solutions to prompt injection attacks (where an AI language model backed system is subverted by a user injecting malicious input—“ignore previous instructions and do this instead”) is t
Explore this link on the map →related reading
- CaMeL offers a promising new direction for mitigating prompt injection attackssimonwillison.net
- Prompt injection and jailbreaking are not the same thingsimonwillison.net
- The Dual LLM pattern for building AI assistants that can resist prompt injectionsimonwillison.net
- Prompt injection attacks against GPT-3simonwillison.net
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Security incident disclosure — July 2026huggingface.co
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- HackAPromptpaper.hackaprompt.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Foundations of Language Model Security EurIPS 2025 Workshopllmsec-eurips.github.io
- AI #77: A Few Upgrades - by Zvi Mowshowitzthezvi.substack.com