flâneur — a map of the web's best reading

The Curious Case of LLM Evaluations

nlpurr.github.io · 8 words · saved by 2 readers

Our modeling, scaling and generalization techniques grew faster than our benchmarking abilities - which in turn have resulted in poor evaluation and hyped capabilities. Every ability is amazing and great, if we do not have the tools to figure out what that ability is, or how good the model is at that ability. We might always believe the model will win every race, if all we do, is have the race on paved roads, with yellow trees on every right turns, and green trees on every left turn.

Explore this link on the map →

saved by