✳flâneur — a map of the web's best reading
HackAPrompt
paper.hackaprompt.com · 783 words · saved by 1 readers
HackAPrompt
HackAPrompt Compete in HackAPrompt 2.0, the world's largest AI Red-Teaming competition! Check it out → × WINNER: Best Theme Paper at EMNLP2023 Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition EMNLP 2023 Sander Schulhoff * 1 , Jeremy Pinto * 2 , Anaum Khan 1 , Louis-François Bouchard 2,3 , Chenglei Si 4 , Svetlina Anati** 5 , Valen Tagliabue** 6 , Anson Liu Kost** 7 , Christopher Carnahan** 8 , Jordan Boyd-Graber 1 ▶ 1 University of Maryland     ▶ 2 Mila     ▶ 3 Towards AI     ▶ 4 Stanford Univ
Explore this link on the map →saved by
related reading
- A small number of samples can poison LLMs of any size \ Anthropicanthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- 2025: The year in LLMssimonwillison.net
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com
- llm-security/README.md at main · greshake/llm-security · GitHubgithub.com
- 2405.01470arxiv.org
- GitHub - SoyGema/pulling_ace · GitHubgithub.com
- Prompts for Work & Play: Launching the Wolfram Prompt Repository-Stephen Wolfram Writingswritings.stephenwolfram.com
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- CaMeL offers a promising new direction for mitigating prompt injection attackssimonwillison.net
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org