flâneur

LLM Evaluation Metrics

mosaicml.com · 220 words · saved by 2 readers

MosaicML's published results of evaluating open-source large language models (LLMs). To evaluate model quality, we compiled 34 open-source benchmarks commonly used for in-context learning (ICL) and evaluated and aggregated them in an industry-standard manner.

saved by