Why I Think More NLP Researchers Should Engage with AI Safety Concerns – NYU Alignment Research Group
Large language modeling research in NLP seems to be feeding into much more impactful technologies than we’re used to working with. While the positive potential for this technology could be tremendous, the downside risk is also potentially catastrophic, and it doesn’t look like we’re prepared to manage that risk. I’m starting a new research group at NYU to work on technical directions that I think are relevant to these concerns, and I’d encourage others in the field to look for ways to take these concerns into account as well. As a research community, we’re making progress quickly but chaotically. The field could nonetheless develop systems that are as good as we are at many important cognitive tasks. If we build AI systems with human-level cognitive abilities, things could get very weird. The only hard requirements for these risks are (i) that systems be competent enough and (ii) that systems be too opaque for us to be able to audit or supervise their reasoning processes reliably. The
Large language modeling research in NLP seems to be feeding into much more impactful technologies than we’re used to working with. While the positive potential for this technology could be tremendous, the downside risk is also potentially catastrophic, and it doesn’t look like we’re prepared to manage that risk. I’m starting a new research group at NYU to work on technical directions that I think are relevant to these concerns, and I’d encourage others in the field to look for ways to take these concerns into account as well. Why I’m concerned As a research community, we’re making progress qui
Explore this link on the map →related reading
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Safety | Arkosevictoriabrook.github.io
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AI safety - Wikipediaen.wikipedia.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- AI in 2025: gestalt — LessWronglesswrong.com
- LLMs for Alignment Research: a safety priority? — AI Alignment Forumalignmentforum.org
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com