Sampling from Your Language Model One Byte at a Time
arxiv.org · 6,405 words · saved by 1 readers
N/A
Sampling from Your Language Model One Byte at a Time Jonathan Hayase 1 Alisa Liu 1 Noah A. Smith 1 2 Sewoong Oh 1 Abstract > olmo.generate(tok.encode("This a tes")) "erstor" Tokenization is used almost universally by mod- >…
saved by
related reading
- The Bitter Lesson is coming for Tokenization – ⛰️ lucalplucalp.dev
- Language Modelinglena-voita.github.io
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4.github.com
- MambaByte: Token-free Selective State Space Modelarxiv.org
- The End of Tokenizationronaldyu.substack.com
- [2112.10508] Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLParxiv.org
- 2409.02908arxiv.org
- Character Prefix Conditioningcursor.com
- Byte Latent Transformer: Patches Scale Better Than Tokensarxiv.org
- Large Language Diffusion Modelsarxiv.org
- Productizing Large Language Modelsblog.replit.com
- I blame the tokenizer | David Quareldavidquarel.github.io