flâneur

GPT-4 and professional benchmarks: the wrong answer to the wrong question

normaltech.ai · 1,933 words · saved by 1 readers

OpenAI may have tested on the training data. Besides, human benchmarks are meaningless for bots.

OpenAI didn’t release much information about GPT-4 — not even the size of the model — but heavily emphasized its performance on professional licensing exams and other standardized tests. For instance, GPT-4 reportedly scored in the 90th percentile on the bar exam. So there’s been much speculation about what this means for professionals such as lawyers. We don’t know the answer, but we hope to inject some reality into the conversation. OpenAI may have violated the cardinal rule of machine learning: don’t test on your training data. Setting that aside, there’s a bigger problem. The manner in…

related reading