flâneur — a map of the web's best reading

Large language models aren’t people. Let’s stop testing them as if they were. | MIT Technology Review

technologyreview.com · 3,310 words · saved by 1 readers

With hopes and fears about this technology running wild, it's time to agree on what it can and can't do.

AI hype is built on high test scores. Those tests are flawed. | MIT Technology Review Skip to Content When Taylor Webb played around with GPT-3 in early 2022, he was blown away by what OpenAI’s large language model appeared to be able to do. Here was a neural network trained only to predict the next word in a block of text —a jumped-up autocomplete. And yet it gave correct answers to many of the abstract problems that Webb set for it—the kind of thing you’d find in an IQ test. “I was really shocked by its ability to solve these problems,” he says. “It completely upended everything I would have

Explore this link on the map →

related reading