Where the goblins came from
Starting with GPT‑5.1, our models began developing a strange habit: they increasingly mentioned goblins, gremlins, and other creatures in their metaphors. Unlike model bugs that show up through a tanking eval or a spiking training metric and point back to a specific change, this one crept in subtly. A single “little goblin” in an answer could be harmless, even charming. Across model generations, though, the habit became hard to miss: the goblins kept multiplying, and we needed to figure out where they came from. In early testing, GPT‑5.5 in Codex showed an odd affinity for goblin metaphors. The short answer is that model behavior is shaped by many small incentives. In this case, one of those incentives came from training the model for the personality customization feature (opens in a new window) , in particular the Nerdy personality. We unknowingly gave particularly high rewards for metaphors with creatures. From there, the goblins spread. The goblins were funny at first, but the incr
April 29, 2026 Publication Where the goblins came from Loading… Share Starting with GPT‑5.1, our models began developing a strange habit: they increasingly mentioned goblins, gremlins, and other creatures in their metaphors. Unlike model bugs that show up through a tanking eval or a spiking training metric and point back to a specific change, this one crept in subtly. A single “little goblin” in an answer could be harmless, even charming. Across model generations, though, the habit became hard to miss: the goblins kept multiplying, and we needed to figure out where they came from. In early tes
Explore this link on the map →saved by
- Yixiong Hao
- Aaron Pham
- Vincent Cheng
- Vincent Huang
- Emil Ryd
- Julian H
- Aadharsh Pannirselvam
- Ishaan Panigrahi
- Anonymous Bat
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- Gwern visits BAIR – Yuxi on the Wiredyuxi.ml
- Expanding on what we missed with sycophancy | OpenAIopenai.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- GPT-4openai.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- gpt-4.pdfcdn.openai.com
- SolidGoldMagikarp (plus, prompt generation) — AI Alignment Forumalignmentforum.org
- Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Diesifanyonebuildsit.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com