Research
As autonomous agents become increasingly woven into the fabric of society—from self-driving cars to personal robot manipulators to AI assistants—our lab aims to ensure their seamless interaction with people. However, integrating these systems into human-centered environments in a way that aligns with human expectations is a formidable challenge. Specifying human objectives to robots is difficult because these objectives are complex, context-dependent, and inherently subjective. Without the right objectives, autonomous systems may exhibit unexpected or even dangerous behaviors. Learning these objectives (for instance, as reward functions) has emerged as a popular alternative to manual specification, but it comes with its own set of difficulties: 1) getting the right data to supervise the learning is hard because humans are imperfect, not infinitely queryable, and have unique and changing preferences; 2) the representations we choose to mathematically express human objectives may themsel
Research CLEAR Interaction @ MIT --> Research Publications People Join As autonomous agents become increasingly woven into the fabric of society—from self-driving cars to personal robot manipulators to AI assistants—our lab aims to ensure their seamless interaction with people. However, integrating these systems into human-centered environments in a way that aligns with human expectations is a formidable challenge. Specifying human objectives to robots is difficult because these objectives are complex, context-dependent, and inherently subjective. Without the right objectives, autonomous syste
Explore this link on the map →related reading
- The Future Worth Building Is Human - Thinking Machines Labthinkingmachines.ai
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The Era of Experience Paper.pdfstorage.googleapis.com
- [AN #70]: Agents that help humans who are still learning about their own preferences — LessWronglesswrong.com
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io
- Learning through human feedback — Google DeepMinddeepmind.google
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Automated Alignment is Harder Than You Think — LessWronglesswrong.com
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- Specification gaming: the flip side of AI ingenuity — Google DeepMinddeepmind.google
- ROGUE:arxiv.org
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com