flâneur

On measuring AI - Nikil Ravi

nikilravi.substack.com · 1,572 words · saved by 1 readers

Thoughts on the state of measurement in AI

Of late, the number of AI model releases per month has been growing very rapidly. Each model’s release is accompanied by the usual benchmark scores that show how well the model performs. Usually, all the charts released by the model provider show their model beating its predecessor and competitors on every reported benchmark. Everything trends up and to the right, this gets reported by various media outlets and reposted by prominent social media accounts, and within a few days, the benchmark scores have either directly or indirectly succeeded in shaping the opinion of most people about the…

saved by

related reading