HackAPrompt
paper.hackaprompt.com · 783 words · saved by 1 readers
HackAPrompt
HackAPrompt Compete in HackAPrompt 2.0, the world's largest AI Red-Teaming competition! Check it out → × WINNER: Best Theme Paper at EMNLP2023 Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition EMNLP 2023 Sander Schulhoff * 1 , Jeremy Pinto * 2 , Anaum Khan 1 , Louis-François Bouchard 2,3 , Chenglei Si 4 , Svetlina Anati** 5 , Valen Tagliabue** 6 , Anson Liu Kost** 7 , Christopher Carnahan** 8 , Jordan Boyd-Graber 1 ▶ 1 University of Maryland     ▶ 2 Mila     ▶ 3 Towards AI     ▶ 4 Stanford Univ
saved by
related reading
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4.github.com
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com
- llm-security/README.md at main · greshake/llm-securitygithub.com
- Prompt Injection as Role Confusionrole-confusion.github.io
- Lakera – Test your AI hacking skillsgandalf.lakera.ai
- Simulated Users & Sad LLMs1a3orn.com
- 2025: The year in LLMssimonwillison.net
- Productizing Large Language Modelsblog.replit.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- CaMeL offers a promising new direction for mitigating prompt injection attackssimonwillison.net
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com