flâneur — a map of the web's best reading

David Gros

dgros.io · 971 words · saved by 1 readers

As machines get more capable at language, it is possible they could deceive humans into thinking they are human. This can be either explicit deception (e.g., a human asks a chatbot “are you a robot?” and it replies incorrectly), or implicit deception (a machine that implies it is human, e.g. “that movie made me cry”). I have researched how to avoid both cases. To explore the explicit case, I worked with collaborators to create the R-U-A-Robot dataset, which collects ~2500 phrasings for how someone might ask if a system is human or non-human. It was published at the ACL 2021 conference. To explore the implicit case, we created the Robots-Dont-Cry dataset which studies kinds of things that are possible for a human to say, but not a machine. It was published at EMNLP 2022. I believe it is useful to work on social and technical progress on "simple norms" (like machines should be honest, and not pretend to be human) as one category of steps to learning about broader AI alignment. Part of my

David Gros Hi, I'm David Gros. I’m a computer scientist and engineer who works in Artificial Intelligence, Machine Learning, and Natural Language Processing. I recently finished a PhD in CS at UC Davis, and previously helped solve interesting problems at Microsoft, IBM Research, and NASA. I primarily focus on how to make AI systems which benefit the world and improving the AI safety of capable systems. Projects --> Here are some of the things I've made in my spare time and a few of the more interesting things I've made for school. --> --> Avoiding machines that pretend to be human As machines

Explore this link on the map →

related reading