Research Publications – Center for Human-Compatible Artificial Intelligence
CHAI aims to reorient the foundations of AI research toward the development of provably beneficial systems. Currently, it is not possible to specify a formula for human values in any form that we know would provably benefit humanity, if that formula were instated as the objective of a powerful AI system. In short, any initial formal specification of human values is bound to be wrong in important ways. This means we need to somehow represent uncertainty in the objectives of AI systems. This way of formulating objectives stands in contrast to the standard model for AI, in which the AI system's objective is assumed to be known completely and correctly. Therefore, much of CHAI's research efforts to date have focussed on developing and communicating a new model of AI development, in which AI systems should be uncertain of their objectives, and should be deferent to humans in light of that uncertainty. However, our interests extend to a variety of other problems in the development of provabl
CHAI aims to reorient the foundations of AI research toward the development of provably beneficial systems . Currently, it is not possible to specify a formula for human values in any form that we know would provably benefit humanity, if that formula were instated as the objective of a powerful AI system. In short, any initial formal specification of human values is bound to be wrong in important ways. This means we need to somehow represent uncertainty in the objectives of AI systems. This way of formulating objectives stands in contrast to the standard model for AI, in which the AI system's
Explore this link on the map →related reading
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org
- Concerns of an Artificial Intelligence Pioneer | Quanta Magazinequantamagazine.org
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- Constitutional AI: Harmlessness from AI Feedbackarxiv.org
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Andi Pengandipeng.com
- Learning through human feedback — Google DeepMinddeepmind.google
- The Iliad Intensive Course Materials — LessWronglesswrong.com