✳flâneur — a map of the web's best reading
TheBloke/Mistral-7B-Instruct-v0.1-AWQ · Hugging Face
huggingface.co · saved by 1 readers
AWQ is an efficient, accurate and blazing-fast low-bit weight quantization method, currently supporting 4-bit quantization. Compared to GPTQ, it offers faster Transformers-based inference.
Explore this link on the map →