Some Lessons from Adversarial Machine Learning | Events at FAR.AI
far.ai · 2,893 words · saved by 1 readers
Some Lessons from Adversarial Machine Learning, 2024 at Vienna Alignment Workshop.
July 20, 2024 Summary SESSION Transcript Great to be here. This is not my usual audience. I'm a computer security person. I'm probably one of the least knowledgeable people about alignment in this room. So maybe what I want to do instead of telling you about what you should be doing, is giving you some words of caution about what might happen in your future if you follow a similar path to what our field in adversarial machine learning followed. Maybe think of this in some sense as a cautionary tale. I am the ghost of Christmas yet to come, telling you what things might look like in ten…
saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Nicholas Carlininicholas.carlini.com
- Solving adversarial attacks in computer vision as a baby version of general AI alignment | Stanislav Fortstanislavfort.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- 2306.15447.pdfarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Narrow Misalignment is Hard, Emergent Misalignment is Easy — LessWronglesswrong.com
- [1610.00768] Technical Report on the CleverHans v2.1.0 Adversarial Examples Libraryarxiv.org
- Adversarial Attacks on Aligned Language Models | Gray Swan Researchgrayswan.ai
- Teaching Claude why \ Anthropicanthropic.com