Evaluation Metrics for Language Modeling
On different metrics for evaluating language models, the relationships among them, mathematical and empirical bounds for those metrics, and suggested best practices with regards to how to report them.
Recently, neural network trained language models, such as ULMFIT, BERT, and GPT-2, have been remarkably successful when transferred to other natural language processing tasks. As such, there's been growing interest in language models. Traditionally, language model performance is measured by perplexity, cross entropy, and bits-per-character (BPC). As language models are increasingly being used as pre-trained models for other NLP tasks, they are often also evaluated based on how well they perform on downstream tasks. The GLUE benchmark score is one example of broader, multi-task evaluation for l
Explore this link on the map →related reading
- Language Modelinglena-voita.github.io
- Successful language model evals - Jason Weijasonwei.net
- [2309.10668] Language Modeling Is Compressionarxiv.org
- Generalized Language Models | Lil'Loglilianweng.github.io
- A History of Large Language Modelsgregorygundersen.com
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- Can gzip be a language model?nathan.rs
- Introduction to Large Language Models | Machine Learning | Google for Developersdevelopers.google.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- arxiv.org/pdf/2505.24832arxiv.org
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Large Language Model: world models or surface statistics?thegradient.pub