EleutherAI/lm-evaluation-harness: A framework for few-shot evaluation of autoregressive language models.
github.com · 5,297 words · saved by 1 readers
A framework for few-shot evaluation of autoregressive language models.
Language Model Evaluation Harness Latest News 📣 [2025/12] CLI refactored with subcommands ( run , ls , validate ) and YAML config file support via --config . See the CLI Reference and Configuration Guide . [2025/12] Lighter install : Base package no longer includes transformers / torch . Install model backends separately: pip install lm_eval[hf] , lm_eval[vllm] , etc. [2025/07] Added think_end_token arg to hf (token/str), vllm and sglang (str) for stripping CoT reasoning traces from models that support it. [2025/03] Added support for steering HF models! [2025/02] Added SGLang support! [2024/0
related reading
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- Hugging Face – The AI community building the future.huggingface.co
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets.github.com
- PostTrainBenchposttrainbench.com
- Successful language model evals - Jason Weijasonwei.net
- The bitter lesson of LLM evalsparsed.com
- Model optimization | OpenAI APIplatform.openai.com
- Localmaxxing - Local LLM Inference Speed Testslocalmaxxing.com
- Goodfire AIgoodfire.ai