Large language models aren’t people. Let’s stop testing them as if they were. | MIT Technology Review
technologyreview.com · 3,310 words · saved by 1 readers
With hopes and fears about this technology running wild, it's time to agree on what it can and can't do.
AI hype is built on high test scores. Those tests are flawed. | MIT Technology Review Skip to Content When Taylor Webb played around with GPT-3 in early 2022, he was blown away by what OpenAI’s large language model appeared to be able to do. Here was a neural network trained only to predict the next word in a block of text —a jumped-up autocomplete. And yet it gave correct answers to many of the abstract problems that Webb set for it—the kind of thing you’d find in an IQ test. “I was really shocked by its ability to solve these problems,” he says. “It completely upended everything I would have
related reading
- gpt-4.pdfcdn.openai.com
- Do large language models understand us? | by Blaise Aguera y Arcas | Mediummedium.com
- How Not to Test GPT-3garymarcus.substack.com
- Large Language Model: world models or surface statistics?thegradient.pub
- [2302.02083] Evaluating Large Language Models in Theory of Mind Tasksarxiv.org
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Cognitive Biases in Large Language Models — LessWronglesswrong.com
- [2302.02083] Theory of Mind May Have Spontaneously Emerged in Large Language Modelsarxiv.org
- The Yale Review | Melanie Mitchell: The Dangerous Unknowns at the…yalereview.org
- Large language models are proficient in solving and creating emotional intelligence testsnature.com
- [2406.02061] Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Modelsarxiv.org
- Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazinequantamagazine.org