flâneur — a map of the web's best reading

I blame the tokenizer | David Quarel

davidquarel.github.io · 1,693 words · saved by 1 readers

To those that developed the tokenizer for the Llama-2 family of language models, I have a bone to pick with you.

To those that developed the tokenizer for the Llama-2 family of language models, I have a bone to pick with you. So I’m procrastinating working on an extension of a paper that claims the Llama-2 family of models “think in English” . They show that if you use logit lens 1 during a forward pass on a multi-shot translation prompt from, say, French to German, internally the model assigns a high probability to the corresponding English word. That’s very curious! I want to query meta-llama/llama-2-7b-hf with a large amount of responses of the form Français: " jour" - Deutsch: " Tag" Français: " homm

Explore this link on the map →

saved by

related reading