Evaluation Metrics for Language Modeling
On different metrics for evaluating language models, the relationships among them, mathematical and empirical bounds for those metrics, and suggested best practices with regards to how to report them.
Recently, neural network trained language models, such as ULMFIT, BERT, and GPT-2, have been remarkably successful when transferred to other natural language processing tasks. As such, there's been growing interest in language models. Traditionally, language model performance is measured by perplexity, cross entropy, and bits-per-character (BPC). As language models are increasingly being used as pre-trained models for other NLP tasks, they are often also evaluated based on how well they perform on downstream tasks. The GLUE benchmark score is one example of broader, multi-task evaluation for l
Explore this link on the map →related reading
- Language Modelinglena-voita.github.io
- [2309.10668] Language Modeling Is Compressionarxiv.org
- Generalized Language Models | Lil'Loglilianweng.github.io
- Can gzip be a language model?nathan.rs
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- arxiv.org/pdf/2505.24832arxiv.org
- Introduction to Large Language Models | Machine Learning | Google for Developersdevelopers.google.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Large Language Model: world models or surface statistics?thegradient.pub
- Language Modeling Without Neural Networksnathan.rs
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- [2001.08361] Scaling Laws for Neural Language Modelsarxiv.org