LLM_Attacks_&_Defenses_Exercise.ipynb - Colab
Note: You can collapse each section so only the headers are visible, by clicking the arrow symbol on the left hand side of the markdown header cells. This workshop is designed to get you familiar with AI jailbreaks. We'll progress from crafting manual jailbreak attacks to implementing automated jailbreak attacks. After each attack method, we'll explore defenses against them. We hope you get a sense of how challenging and rewarding the cat-and-mouse game of AI security is! Lastly, we'll cover more advanced methods theoretically and have links to further resources. Language models such as ChatGPT, Gemini and Claude have seen widespread deployment due to their advanced capabilities. However, they are also suspect to misuse by bad actors. To combat this, researchers have implemented safety mechanisms such as aligning model behaviors with human feedback. While these alignment techniques help, people soon found out that models are susceptible to jailbreaking—carefully crafted prompts that ta
Explore this link on the map →