LLM.int8() and Emergent Features — Tim Dettmers
When I attended NAACL, I wanted to do a little test. I had two pitches for my LLM.int8() paper. One pitch is about how I use advanced quantization methods to achieve no performance degradation transformer inference at scale that makes large models more accessible. The other pitch talks about emergent outliers in transformers and how […]
When I attended NAACL, I wanted to do a little test. I had two pitches for my LLM.int8() paper. One pitch is about how I use advanced quantization methods to achieve no performance degradation transformer inference at scale that makes large models more accessible. The other pitch talks about emergent outliers in transformers and how they radically change what transformers learn and how they function. From that, I learned that quantization research is like printers. Nobody cares about printers. Nobody likes printers. But everybody is happy if printers do their job. How that job is done for you
Explore this link on the map →related reading
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- LLM.int8()arxiv.org
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnettransformer-circuits.pub
- Attention Is Off By One – Evan Millerevanmiller.org
- Transformer Circuits Threadtransformer-circuits.pub
- Transformers from Scratche2eml.school
- The Annotated Transformernlp.seas.harvard.edu
- On neural scaling and the quanta hypothesisericjmichaud.com
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- [2304.07493] OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantizationarxiv.org
- Transformer Math 101 | EleutherAI Blogblog.eleuther.ai