Perplexity is not a good measurement of how well a model actually performs : r/LocalLLaMA
Subreddit to discuss AI & Llama, the large language model created by Meta AI. Hey, Tony I'm making this post because I've been on this sub since the early llama 1 days and I've seen perplexity used as a sort of de-facto "gold standard" when it comes to talking about quantization methods and model performance. Regularly there will be posts here along the lines of "new N-bit quantization technique achieves minimal perplexity loss" or something along those lines. To be clear I'm nowhere near an expert when it comes to LLMs and this post is absolutely not trying to shit on any of the many awesome advancements that we are seeing daily in the open source LLM community. I just wanted to start this discussion so more people can become aware of the limitations of perplexity as a proxy for model performance and maybe try discuss some alternative ways of evaluating models. I've spent lots of time experimenting with different fine-tunes and various quantization techniques. I remember using GPTQ an
Explore this link on the map →