Language models are weird for the same reason human cultures are weird
substack.com · 4,360 words · saved by 1 readers
You can’t have adaptive learning without strange tics
In November 2025, shortly after OpenAI released GPT-5.1—a new model that promised “a smarter, more conversational ChatGPT”—a small set of users started to notice something weird. GPT-5.1 was indeed smarter and more conversational; but it also had a strange habit of referring to things as “goblins.” For a time this was treated as a quirk—language models, after all, do all sorts of strange things—and nobody gave it much thought. But soon things started to get stranger. With each new model release in the months that followed—5.2 in December, 5.3 in February, 5.4 in March, 5.5 in April—OpenAI’s…
related reading
- Where the goblins came from | OpenAIopenai.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Gwern visits BAIR – Yuxi on the Wiredyuxi.ml
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Towards a Typology of Strange LLM Chains-of-Thought1a3orn.com
- Late Takes on OpenAI o1alexirpan.com
- How AI Is Learning to Think in Secretnickandresen.substack.com
- SolidGoldMagikarp (plus, prompt generation) — AI Alignment Forumalignmentforum.org
- Simulated Users & Sad LLMs1a3orn.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com