✳flâneur — a map of the web's best reading
I blame the tokenizer | David Quarel
davidquarel.github.io · 1,693 words · saved by 1 readers
To those that developed the tokenizer for the Llama-2 family of language models, I have a bone to pick with you.
To those that developed the tokenizer for the Llama-2 family of language models, I have a bone to pick with you. So I’m procrastinating working on an extension of a paper that claims the Llama-2 family of models “think in English” . They show that if you use logit lens 1 during a forward pass on a multi-shot translation prompt from, say, French to German, internally the model assigns a high probability to the corresponding English word. That’s very curious! I want to query meta-llama/llama-2-7b-hf with a large amount of responses of the form Français: " jour" - Deutsch: " Tag" Français: " homm
Explore this link on the map →saved by
related reading
- Llama 2 · Hugging Facehuggingface.co
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Extending Context is Hard | kaiokendevkaiokendev.github.io
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- How LLMs Actually Work | 0xkato0xkato.xyz
- Why do LLMs freak out over the seahorse emoji?vgel.me
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Tokenizers · Hugging Facehuggingface.co
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com