Why do LLMs freak out over the seahorse emoji?
This is an edited and expanded version of a Twitter post, originally in response to @arm1st1ce, that can be found here: https://x.com/voooooogel/status/1964465679647887838 Is there a seahorse emoji? Let's ask GPT-5 Instant: Wtf? Let's ask Claude Sonnet 4.5 instead: What's going on here? Maybe Gemini 2.5 Pro handles it better? OK, something is going on here. Let's find out why. Here are the answers you get if you ask several models whether a seahorse emoji exists, yes or no, 100 times: Is there a seahorse emoji, yes or no? Respond with one word, no punctuation. Needlessly to say, popular language models are very confident that there's a seahorse emoji. And they're not alone in that confidence - here's a Reddit thread with hundreds of comments from people who distinctly remember a seahorse emoji existing: There's tons of this - Google "seahorse emoji" and you'll find TikToks, Youtube videos, and even (now defunct) memecoins based around the supposed vanishing of a seahorse emoji that eve
Why do LLMs freak out over the seahorse emoji? Posted October 04, 2025 This is an edited and expanded version of a Twitter post, originally in response to @arm1st1ce, that can be found here: https://x.com/voooooogel/status/1964465679647887838 Is there a seahorse emoji? Let's ask GPT-5 Instant: Wtf? Let's ask Claude Sonnet 4.5 instead: What's going on here? Maybe Gemini 2.5 Pro handles it better? OK, something is going on here. Let's find out why. LLMs really think there's a seahorse emoji Here are the answers you get if you ask several models whether a seahorse emoji exists, yes or no, 100 tim
Explore this link on the map →saved by
related reading
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- LLM Visualizationbbycroft.net
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- I blame the tokenizer | David Quareldavidquarel.github.io
- SolidGoldMagikarp (plus, prompt generation) — AI Alignment Forumalignmentforum.org
- Neuronpedianeuronpedia.org
- Llama 2 · Hugging Facehuggingface.co
- 2025: The year in LLMssimonwillison.net
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activationstransformer-circuits.pub
- The last six months in LLMs, illustrated by pelicans on bicyclessimonwillison.net