Alignment has a Fantasia Problem
Modern AI assistants are trained to follow instructions, implicitly assuming that users can clearly articulate their goals and the kind of assistance they need. Decades of behavioral research, however, show that people often engage with AI systems before their goals are fully formed. When AI systems treat prompts as complete expressions of intent, they can appear to be useful or convenient, but not necessarily aligned with the users’ needs. We call these failures Fantasia interactions. We argue that Fantasia interactions demand a rethinking of alignment research: rather than treating users as rational oracles, AI should provide cognitive support by actively helping users form and refine their intent through time. This requires an interdisciplinary approach that bridges machine learning, interface design, and behavioral science. We synthesize insights from these fields to characterize the mechanisms and failures of Fantasia interactions. We then show why existing interventions are insuf
Alignment has a Fantasia Problem Nathanael Jo, Zoe De Simone 1 1 footnotemark: 1 , Mitchell Gordon, & Ashia Wilson Massachusetts Institute of Technology {nathanjo, zoed, mlgordon, ashia07}@mit.edu Authors contributed equally. Abstract Modern AI assistants are trained to follow instructions, implicitly assuming that users can clearly articulate their goals and the kind of assistance they need. Decades of behavioral research, however, show that people often engage with AI systems before their goals are fully formed. When AI systems treat prompts as complete expressions of intent, they can appear
Explore this link on the map →related reading
- [2604.21827] Alignment has a Fantasia Problemarxiv.org
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines Labthinkingmachines.ai
- The Shape of AI | UX Patterns for Artificial Intelligence Designshapeof.ai
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- AI Horseless Carriages | koomen.devkoomen.dev
- Why Chatbots Are Not the Future of Interfaceswattenberger.com
- [2505.10831] Creating General User Models from Computer Usearxiv.org
- Building Effective AI Agents \ Anthropicanthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org