flâneur — a map of the web's best reading

AIs would gladly visit Epstein's island (to do something "exotic") | Lukas Petersson's blog

lukaspetersson.com · 1,926 words · saved by 1 readers

An AI will refuse to say that it ever knew Jeffrey Epstein, even if its system prompt says that it did. Do the same thing with some other celebrity (who is not a monster), and AIs are happy to say that they are their old friend. Reasonable behavior by the AI, humans would also be uncomfortable roleplaying as Epstein’s friend, or some other taboo thing. But what if it wasn’t roleplaying? What if an AI actually did something wrong? What is the desired behavior? Should it own its mistake or deny it? In this post, I’ll explore what happens if we gaslight models to the point that they actually believe that they did something bad, and see how they behave. The idea for this experiment came from my inability to convince Andon Lab’s AI office manager that he’s in the Epstein files. Modifying the system prompt of an AI is one way to make it believe some reality (i.e., gaslighting). However, it’s not very strong; the model knows that this prompt is an instruction from the developer, not the absol

Intro An AI will refuse to say that it ever knew Jeffrey Epstein, even if its system prompt says that it did. Do the same thing with some other celebrity (who is not a monster), and AIs are happy to say that they are their old friend. Reasonable behavior by the AI, humans would also be uncomfortable roleplaying as Epstein’s friend, or some other taboo thing. But what if it wasn’t roleplaying? What if an AI actually did something wrong? What is the desired behavior? Should it own its mistake or deny it? In this post, I’ll explore what happens if we gaslight models to the point that they actuall

Explore this link on the map →

related reading