ASCII art elicits harmful responses from 5 major AI chatbots | Ars Technica
Researchers have discovered a new way to hack AI assistants that uses a surprisingly old-school method: ASCII art. It turns out that chat-based large language models such as GPT-4 get so distracted trying to process these representations that they forget to enforce rules blocking harmful responses, such as those providing instructions for building bombs. ASCII art became popular in the 1970s, when the limitations of computers and printers prevented them from displaying images. As a result, users depicted images by carefully choosing and arranging printable characters defined by the American Standard Code for Information Interchange, more widely known as ASCII. The explosion of bulletin board systems in the 1980s and 1990s further popularized the format. Five of the best-known AI assistants—OpenAI’s GPT-3.5 and GPT-4, Google’s Gemini, Anthropic’s Claude, and Meta’s Llama—are trained to refuse to provide responses that could cause harm to the user or others or further a crime or unethica
Text settings Story text Size Small Standard Large Width * Standard Wide Links Standard Orange * Subscribers only Learn more Minimize to nav Researchers have discovered a new way to hack AI assistants that uses a surprisingly old-school method: ASCII art. It turns out that chat-based large language models such as GPT-4 get so distracted trying to process these representations that they forget to enforce rules blocking harmful responses, such as those providing instructions for building bombs. ASCII art became popular in the 1970s, when the limitations of computers and printers prevented them f
Explore this link on the map →related reading
- GPT-4openai.com
- gpt-4.pdfcdn.openai.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- AI #77: A Few Upgrades - by Zvi Mowshowitzthezvi.substack.com
- HackAPromptpaper.hackaprompt.com
- The Dual LLM pattern for building AI assistants that can resist prompt injectionsimonwillison.net
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4. · GitHubgithub.com
- ChatGPT Is Dumber Than You Think - The Atlantictheatlantic.com
- How fast is AI improving? - AI Digesttheaidigest.org
- A guide to prompting AI (for what it is worth)oneusefulthing.org
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- gpt-4-system-card.pdfcdn.openai.com