Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Dies
ifanyonebuildsit.com · 2,721 words · saved by 1 readers
Resources and Q&A for the book If Anyone Builds It, Everyone Dies.
Aren’t developers regularly making their AIs nice and safe and obedient? AIs steer in alien directions that only mostly coincide with helpfulness. Modern AIs are pretty helpful (or at least not harmful) to most users, most of the time. But as we noted above , a critical question is how to distinguish an AI that deeply wants to be helpful and do the right thing, from an AI with weirder and more complex drives that happen to line up with helpfulness under typical conditions, but which would prefer other conditions and outcomes even more. * Both sorts of AIs would act helpful in the typical case.
saved by
related reading
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Unfalsifiable stories of doom | Mechanize, Inc.mechanize.work
- Teaching Claude Whyalignment.anthropic.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- AI #87: Staying in Character - by Zvi Mowshowitzthezvi.substack.com
- Claude’s Character \ Anthropicanthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com