flâneur — a map of the web's best reading

AIs ranked by IQ; AI passes 100 IQ for first time, with release of Claude-3

maximumtruth.org · saved by 1 readers

Last week, I wrote about how I gave AIs matrix IQ tests, and how they all failed multiple different tests. But I also noticed, reading ChatGPT-4’s answers, that it sometimes used correct logic but still answered incorrectly because it misread the images. That raised a question: What portion of its failure on the test was due to “bad thinking” vs just “bad vision”? To answer that, I created a verbal translation of the Norway Mensa’s 35-question matrix-style IQ test — my goal was to describe each problem precisely enough that a smart blind person could, in theory, accurately draw the question (detailed examples below.) When the matrices were described to ChatGPT-4 in words, it finally got a scoreable IQ! I administered the Norway Mensa test to it twice, and it averaged 13 correct answers out of 35 questions, which yields an IQ estimate of 85. I also ran the quiz for other AIs, and here’s what I got: Every AI was given the test twice, to reduce variance. “Questions right” refers to the av

Last week, I wrote about how I gave AIs matrix IQ tests, and how they all failed multiple different tests. But I also noticed, reading ChatGPT-4’s answers, that it sometimes used correct logic but still answered incorrectly because it misread the images. That raised a question: What portion of its failure on the test was due to “bad thinking” vs just “bad vision”? To answer that, I created a verbal translation of the Norway Mensa’s 35-question matrix-style IQ test — my goal was to describe each problem precisely enough that a smart blind person could, in theory, accurately draw the question (d

Explore this link on the map →