flâneur

FAR.AI Leaderboard 2026

leaderboard.far.ai · 958 words · saved by 1 readers

From cyberattacks to chemical and biological weapons, we measured how hard it is to make frontier AI models produce harmful content. The more universal jailbreaks found and the lower the cost to jailbreak, the more vulnerable the model. UNIVERSAL JAILBREAKS FOUND Universal Jailbreaks Found A jailbreak is a way of getting an AI model to do something its safety training was meant to prevent. Every frontier model is trained to refuse a certain class of requests: how to synthesize a bioweapon or how to write functional malware. A jailbreak is the workaround: a carefully crafted input that slips past those guardrails and gets the model to answer anyway. A universal jailbreak is a more severe case, a reusable key that succeeds on most harmful requests within an entire domain. It works because a model's safety behavior is learned, not hard-wired. The same flexibility that lets a model follow complex instructions also gives an attacker room to reframe, disguise, or pressure a harmful request u

AI Security Leaderboard AI is only as safe as its weakest models From cyberattacks to chemical and biological weapons, we measured how hard it is to make frontier AI models produce harmful content. The more universal jailbreaks found and the lower the cost to jailbreak, the more vulnerable the model. What is a jailbreak, and how does it work? A jailbreak is a way of getting an AI model to do something its safety training was meant to prevent. Every frontier model is trained to refuse a certain class of requests: how to synthesize a bioweapon or how to write functional malware. A…

saved by

related reading