FAR.AI Leaderboard 2026
From cyberattacks to chemical and biological weapons, we measured how hard it is to make frontier AI models produce harmful content. The more universal jailbreaks found and the lower the cost to jailbreak, the more vulnerable the model. UNIVERSAL JAILBREAKS FOUND Universal Jailbreaks Found A jailbreak is a way of getting an AI model to do something its safety training was meant to prevent. Every frontier model is trained to refuse a certain class of requests: how to synthesize a bioweapon or how to write functional malware. A jailbreak is the workaround: a carefully crafted input that slips past those guardrails and gets the model to answer anyway. A universal jailbreak is a more severe case, a reusable key that succeeds on most harmful requests within an entire domain. It works because a model's safety behavior is learned, not hard-wired. The same flexibility that lets a model follow complex instructions also gives an attacker room to reframe, disguise, or pressure a harmful request u
AI Security Leaderboard AI is only as safe as its weakest models From cyberattacks to chemical and biological weapons, we measured how hard it is to make frontier AI models produce harmful content. The more universal jailbreaks found and the lower the cost to jailbreak, the more vulnerable the model. What is a jailbreak, and how does it work? A jailbreak is a way of getting an AI model to do something its safety training was meant to prevent. Every frontier model is trained to refuse a certain class of requests: how to synthesize a bioweapon or how to write functional malware. A…
saved by
related reading
- Seoul Alignment Workshop 2026: What We Learnedfar.ai
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- The End-State Fallacy: Where Is AI Security Headed?endstatefallacy.com
- A Safe Path to Open Weights - Thinking Machines Labthinkingmachines.ai
- Common Elements of Frontier AI Safety Policiesmetr.org
- CAIS AI Dashboarddashboard.safe.ai
- Redeploying Claude Fable 5 \ Anthropicanthropic.com
- Automatically Jailbreaking Frontier Language Models with Investigator Agents | Transluce AItransluce.org
- [2602.15001] Boundary Point Jailbreaking of Black-Box LLMsarxiv.org
- Cost-Effective Constitutional Classifiers via Representation Re-usealignment.anthropic.com
- Security incident disclosure — July 2026huggingface.co
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaksarxiv.org