Evaluation Metrics for Language Modeling
On different metrics for evaluating language models, the relationships among them, mathematical and empirical bounds for those metrics, and suggested best practices with regards to how to report them.
Recently, neural network trained language models, such as ULMFIT, BERT, and GPT-2, have been remarkably successful when transferred to other natural language processing tasks. As such, there's been growing interest in language models. Traditionally, language model performance is measured by perplexity, cross entropy, and bits-per-character (BPC). As language models are increasingly being used as pre-trained models for other NLP tasks, they are often also evaluated based on how well they perform on downstream tasks. The GLUE benchmark score is one example of broader, multi-task evaluation for l
related reading
- Language Modelinglena-voita.github.io
- Successful language model evals - Jason Weijasonwei.net
- [2309.10668] Language Modeling Is Compressionarxiv.org
- Compression and Intelligencegreene.sh
- Generalized Language Models | Lil'Loglilianweng.github.io
- [2005.14165] Language Models are Few-Shot Learnersarxiv.org
- A History of Large Language Modelsgregorygundersen.com
- Can gzip be a language model?nathan.rs
- Introduction to Large Language Models | Machine Learning | Google for Developersdevelopers.google.com
- [2411.00640] Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluationsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- arxiv.org/pdf/2505.24832arxiv.org