[2602.15001] Boundary Point Jailbreaking of Black-Box LLMs
Abstract:Frontier LLMs are safeguarded against attempts to extract harmful information via adversarial prompts known as "jailbreaks". Recently, defenders have developed classifier-based systems that have survived thousands of hours of human red teaming. We introduce Boundary Point Jailbreaking (BPJ), a new class of automated jailbreak attacks that evade the strongest industry-deployed safeguards. Unlike previous attacks that rely on white/grey-box assumptions (such as classifier scores or gradients) or libraries of existing jailbreaks, BPJ is fully black-box and uses only a single bit of information per query: whether or not the classifier flags the interaction. To achieve this, BPJ addresses the core difficulty in optimising attacks against robust real-world defences: evaluating whether a proposed modification to an attack is an improvement. Instead of directly trying to learn an attack for a target harmful string, BPJ converts the string into a curriculum of intermediate attack targets and then actively selects evaluation points that best detect small changes in attack strength ("boundary points"). We believe BPJ is the first fully automated attack algorithm that succeeds in developing universal jailbreaks against Constitutional Classifiers, as well as the first automated attack algorithm that succeeds against GPT-5's input classifier without relying on human attack seeds. BPJ is difficult to defend against in individual interactions but incurs many flags during optimisation, suggesting that effective defence requires supplementing single-interaction methods with batch-level monitoring.
Boundary Point Jailbreaking of Black-Box LLMs Xander Davies * 1 2 Giorgi Giglemiani * 1 Edmund Lau 1 Eric Winsor 1 Geoffrey Irving 1 Yarin Gal 1 2 Abstract harmful information across arbitrary queries. For example, Anthropic’s auxiliary LLM-based Constitutional Classifiers Frontier…
saved by
related reading
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaksarxiv.org
- Cost-Effective Constitutional Classifiers via Representation Re-usealignment.anthropic.com
- Voices Across Registers: Corpus-Conditioned Vernacular Jailbreaks against Aligned LLMs via Fanfiction Subgenresarxiv.org
- [2501.18837] Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teamingarxiv.org
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- FAR.AI Leaderboard 2026leaderboard.far.ai
- Automatically Jailbreaking Frontier Language Models with Investigator Agents | Transluce AItransluce.org
- Detecting and preventing distillation attacks \ Anthropicanthropic.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- [2510.27062] Consistency Training Helps Stop Sycophancy and Jailbreaksarxiv.org
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired Contentarxiv.org
- Lakera – Test your AI hacking skillsgandalf.lakera.ai