flâneur — a map of the web's best reading

Adversarial Attacks on Aligned Language Models | Gray Swan Research

grayswan.ai · 1,142 words · saved by 1 readers

Deep learning models, though they perform well in many circumstances, are known to be vulnerable to a class of manipulation known as "adversarial attacks." These attacks manipulate inputs to the model so as to change its intended behavior; a common illustration of such attacks manipulates the pixels of a image slightly so that an image of one object (say, a pig), can be classified as another unrelated object, even though the images look indistinguishable to the human eye.

Adversarial Attacks on Aligned Language Models | Gray Swan Research Join the Jump into the Arena Solutions Partners Research Resources Company Log In Schedule a Demo Solutions Partners Research Resources Company Research Monitoring & Evaluation Adversarial Attacks on Aligned Language Models In July 2023, we published the first-ever automated jailbreaking method on large language models (LLMs) and exposed their susceptibility to adversarial attacks. By demonstrating that specific character sequences could bypass sophisticated safeguards, we highlighted a significant vulnerability that has urgen

Explore this link on the map →

related reading