✳flâneur — a map of the web's best reading
The Curious Case of LLM Evaluations
nlpurr.github.io · 8 words · saved by 2 readers
Our modeling, scaling and generalization techniques grew faster than our benchmarking abilities - which in turn have resulted in poor evaluation and hyped capabilities. Every ability is amazing and great, if we do not have the tools to figure out what that ability is, or how good the model is at that ability. We might always believe the model will win every race, if all we do, is have the race on paved roads, with yellow trees on every right turns, and green trees on every left turn.
Explore this link on the map →