GPT-4 and professional benchmarks: the wrong answer to the wrong question
normaltech.ai · 1,933 words · saved by 1 readers
OpenAI may have tested on the training data. Besides, human benchmarks are meaningless for bots.
OpenAI didn’t release much information about GPT-4 — not even the size of the model — but heavily emphasized its performance on professional licensing exams and other standardized tests. For instance, GPT-4 reportedly scored in the 90th percentile on the bar exam. So there’s been much speculation about what this means for professionals such as lawyers. We don’t know the answer, but we hope to inject some reality into the conversation. OpenAI may have violated the cardinal rule of machine learning: don’t test on your training data. Setting that aside, there’s a bigger problem. The manner in…
related reading
- gpt-4.pdfcdn.openai.com
- GPT-4openai.com
- 2303.08774.pdfarxiv.org
- I. From GPT-4 to AGI: Counting the OOMs - SITUATIONAL AWARENESSsituational-awareness.ai
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- PostTrainBenchposttrainbench.com
- gpt-4-system-card.pdfcdn.openai.com
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4.github.com
- What will GPT-2030 look like? — AI Alignment Forumalignmentforum.org
- Every Benchmark is Brokenjonathanpgabor.substack.com
- 2025: The year in LLMssimonwillison.net
- AI hype is built on high test scores. Those tests are flawed. | MIT Technology Reviewtechnologyreview.com