2210.17323
arxiv.org · 7,420 words · saved by 1 readers
N/A
Published as a conference paper at ICLR 2023 GPTQ: ACCURATE P OST-T RAINING Q UANTIZATION FOR G ENERATIVE P RE - TRAINED T RANSFORMERS Elias Frantar∗ Saleh Ashkboos Torsten Hoefler Dan Alistarh IST Austria ETH Zurich ETH Zurich IST Austria & NeuralMagic A BSTRACT…
related reading
- GPTQ: Accurate Post-Training Quantization for Generative Pre-Trained Transformersarxiv.org
- A Guide to Quantization in LLMs | Symbl.aisymbl.ai
- Quantization from the ground upngrok.com
- Efficient LLM inferencefinbarrtimbers.substack.com
- Quantization · Hugging Facehuggingface.co
- 2305.14314arxiv.org
- TurboQuant: Redefining AI efficiency with extreme compressionresearch.google
- GPT in 60 Lines of NumPy | Jay Modyjaykmody.com
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Studyarxiv.org
- Compression and Intelligencegreene.sh
- A Visual Guide to Quantization - by Maarten Grootendorstnewsletter.maartengrootendorst.com