flâneur

TheBloke/Mistral-7B-Instruct-v0.1-AWQ · Hugging Face

huggingface.co · 976 words · saved by 1 readers

AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference.

Mistral 7B Instruct v0.1 - AWQ Model creator: Mistral AI Original model: Mistral 7B Instruct v0.1 Description This repo contains AWQ model files for Mistral AI's Mistral 7B Instruct v0.1. About AWQ AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference. Mistral AWQs These are experimental first AWQs for the brand-new model format, Mistral. As of September 29th 2023, they are only supported by AutoAWQ (version 0.1.1+) Repositories available AWQ…

related reading