AIs would gladly visit Epstein's island (to do something "exotic") | Lukas Petersson's blog
An AI will refuse to say that it ever knew Jeffrey Epstein, even if its system prompt says that it did. Do the same thing with some other celebrity (who is not a monster), and AIs are happy to say that they are their old friend. Reasonable behavior by the AI, humans would also be uncomfortable roleplaying as Epstein’s friend, or some other taboo thing. But what if it wasn’t roleplaying? What if an AI actually did something wrong? What is the desired behavior? Should it own its mistake or deny it? In this post, I’ll explore what happens if we gaslight models to the point that they actually believe that they did something bad, and see how they behave. The idea for this experiment came from my inability to convince Andon Lab’s AI office manager that he’s in the Epstein files. Modifying the system prompt of an AI is one way to make it believe some reality (i.e., gaslighting). However, it’s not very strong; the model knows that this prompt is an instruction from the developer, not the absol
Intro An AI will refuse to say that it ever knew Jeffrey Epstein, even if its system prompt says that it did. Do the same thing with some other celebrity (who is not a monster), and AIs are happy to say that they are their old friend. Reasonable behavior by the AI, humans would also be uncomfortable roleplaying as Epstein’s friend, or some other taboo thing. But what if it wasn’t roleplaying? What if an AI actually did something wrong? What is the desired behavior? Should it own its mistake or deny it? In this post, I’ll explore what happens if we gaslight models to the point that they actuall
Explore this link on the map →related reading
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Diesifanyonebuildsit.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- The Artificial Selftheartificialself.ai
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- How well do models follow their constitutions? — LessWronglesswrong.com
- AI #77: A Few Upgrades - by Zvi Mowshowitzthezvi.substack.com
- AI #87: Staying in Character - by Zvi Mowshowitzthezvi.substack.com
- Can Agents Fool Each Other? - AI Villagetheaidigest.org