David Gros
As machines get more capable at language, it is possible they could deceive humans into thinking they are human. This can be either explicit deception (e.g., a human asks a chatbot “are you a robot?” and it replies incorrectly), or implicit deception (a machine that implies it is human, e.g. “that movie made me cry”). I have researched how to avoid both cases. To explore the explicit case, I worked with collaborators to create the R-U-A-Robot dataset, which collects ~2500 phrasings for how someone might ask if a system is human or non-human. It was published at the ACL 2021 conference. To explore the implicit case, we created the Robots-Dont-Cry dataset which studies kinds of things that are possible for a human to say, but not a machine. It was published at EMNLP 2022. I believe it is useful to work on social and technical progress on "simple norms" (like machines should be honest, and not pretend to be human) as one category of steps to learning about broader AI alignment. Part of my
David Gros Hi, I'm David Gros. I’m a computer scientist and engineer who works in Artificial Intelligence, Machine Learning, and Natural Language Processing. I recently finished a PhD in CS at UC Davis, and previously helped solve interesting problems at Microsoft, IBM Research, and NASA. I primarily focus on how to make AI systems which benefit the world and improving the AI safety of capable systems. Projects --> Here are some of the things I've made in my spare time and a few of the more interesting things I've made for school. --> --> Avoiding machines that pretend to be human As machines
Explore this link on the map →related reading
- ChatGPT Is Nothing Like a Human, Says Linguist Emily Bendernymag.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- The Dark Forest and Generative AImaggieappleton.com
- Import AI 431: Technological Optimism and Appropriate Fearimportai.substack.com
- AI Safety Seems Hard to Measurecold-takes.com
- Import AIjack-clark.net
- How confessions can keep language models honest | OpenAIopenai.com
- AI #77: A Few Upgrades - by Zvi Mowshowitzthezvi.substack.com
- The Yale Review | Melanie Mitchell: The Dangerous Unknowns at the…yalereview.org
- What Is Claude? Anthropic Doesn’t Know, Either | The New Yorkernewyorker.com
- LLMs Go To Confession, Automated Scientific Research, What Copilot Users Want, and more...deeplearning.ai