flâneur

2306.15447.pdf

arxiv.org · 8,651 words · saved by 1 readers

N/A

Are aligned neural networks adversarially aligned? Nicholas Carlini1 , Milad Nasr1 , Christopher A. Choquette-Choo1 , Matthew Jagielski1 , Irena Gao2 , Anas Awadalla3 , Pang Wei Koh13 , Daphne Ippolito1 , Katherine Lee1 , Florian Tramèr4 , Ludwig Schmidt3 1 Google DeepMind 2 Stanford 3 University of Washington 4 ETH Zurich…

related reading