flâneur

Indirect Prompt Injection Attacks LLMs

github.com · 1,127 words · saved by 1 readers

New ways of breaking app-integrated LLMs . Contribute to greshake/llm-security development by creating an account on GitHub.

New: Demonstrating Indirect Injection attacks on Bing Chat Compromising LLMs using Indirect Prompt Injection "... a language model is a Turing-complete weird machine running programs written in natural language; when you do retrieval, you are not 'plugging updated facts into your AI', you are actually downloading random new unsigned blobs of code from the Internet (many written by adversaries) and casually executing them on your LM with full privileges. This does not end well." - Gwern Branwen on LessWrong We present a new class of vulnerabilities and impacts stemming from "indirect…

saved by

related reading