✳flâneur — a map of the web's best reading
EleutherAI/lm-evaluation-harness: A framework for few-shot evaluation of autoregressive language models.
github.com · 5,297 words · saved by 1 readers
A framework for few-shot evaluation of autoregressive language models.
Language Model Evaluation Harness Latest News 📣 [2025/12] CLI refactored with subcommands ( run , ls , validate ) and YAML config file support via --config . See the CLI Reference and Configuration Guide . [2025/12] Lighter install : Base package no longer includes transformers / torch . Install model backends separately: pip install lm_eval[hf] , lm_eval[vllm] , etc. [2025/07] Added think_end_token arg to hf (token/str), vllm and sglang (str) for stripping CoT reasoning traces from models that support it. [2025/03] Added support for steering HF models! [2025/02] Added SGLang support! [2024/0
Explore this link on the map →related reading
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de
- The bitter lesson of LLM evalsparsed.com
- Model optimization | OpenAI APIplatform.openai.com
- PostTrainBenchposttrainbench.com
- GitHub - openai/parameter-golf: Train the smallest LM you can that fits in 16MB. Best model wins! · GitHubgithub.com
- GitHub - karpathy/nanochat: The best ChatGPT that $100 can buy. · GitHubgithub.com
- Neuronpedianeuronpedia.org
- GitHub - SoyGema/pulling_ace · GitHubgithub.com
- GitHub - open-compass/opencompass: OpenCompass is an LLM evaluation platform, supporting a wide range of models (Llama3, Mistral, InternLM2,GPT-4,LLaMa2, Qwen,GLM, Claude, etc) over 100+ datasets. · GitHubgithub.com
- Branches · HazyResearch/intelligence-per-watt · GitHubgithub.com