Vocab Break – Ian’s Blog
Tokenizer enthusiast Sander Land recently reproduced something very like Claude’s current tokenizer, and it appears to only have about 16,000 entries. That is surprising! Qwen 3.8, a very strong re…
Tokenizer enthusiast Sander Land recently reproduced something very like Claude’s current tokenizer, and it appears to only have about 16,000 entries. That is surprising! Qwen 3.8, a very strong release, has about 250k tokens in its vocab. In general the trend had seemed to be more is better in this space. One theory is that Anthropic have been working around a bottleneck caused by the final softmax layer. There is a recent(ish) paper about this: “Lost in Backpropagation: The LM Head is a Gradient Bottleneck“, but, if this is the reason, then the folks at Throppy have known this for way…
saved by
related reading
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- The End of Tokenizationronaldyu.substack.com
- LLMs can invent their own compression - Rajan Agarwalrajan.sh
- microgptkarpathy.github.io
- The Annotated Transformernlp.seas.harvard.edu
- H-Nets - the Past | Goomba Labgoombalab.github.io
- You Can Learn Tokenization End-to-End with Reinforcement Learningarxiv.org
- ali (@waterloo_intern) on Xx.com
- The Big LLM Architecture Comparisonmagazine.sebastianraschka.com