'The illusion of thinking': Apple research finds AI models collapse and give up with hard puzzles | Mashable
New artificial intelligence research from Apple shows AI reasoning models may not be "thinking" so well after all. According to a paper published just days before Apple's WWDC event, large reasoning models (LRMs) — like OpenAI o1 and o3, DeepSeek R1, Claude 3.7 Sonnet Thinking, and Google Gemini Flash Thinking — completely collapse when they're faced with increasingly complex problems. The paper comes from the same researchers who found other reasoning flaws in LLMs last year. The news was a bucket of cold water for artificial general intelligence (AGI) optimists (and welcome news for AI and AGI skeptics), as Apple's research seemed to show damning evidence about the limitations of reasoning model intelligence. While the much-hyped LRM performed better than LLMs on medium-difficulty puzzles, they performed worse on simple puzzles. And according to Apple's research, when they faced hard puzzles, they collapsed completely, giving up on the problem prematurely. Or, as the Apple researcher
The Tower of Hanoi puzzle is too much for reasoning models at a certain point. Credit: CorbalanStudio / iStock / Getty Images New artificial intelligence research from Apple shows AI reasoning models may not be "thinking" so well after all. According to a paper published just days before Apple's WWDC event , large reasoning models (LRMs) - like OpenAI o1 and o3, DeepSeek R1, Claude 3.7 Sonnet Thinking, and Google Gemini Flash Thinking - completely collapse when they're faced with increasingly complex problems. The paper comes from the same researchers who found other reasoning flaws in LLMs la
Explore this link on the map →related reading
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- DeepSeek-R1arxiv.org
- Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazinequantamagazine.org
- [2406.02061] Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Modelsarxiv.org
- Reasoning Models Reason Well, Until They Don'tarxiv.org
- A knockout blow for LLMs? - by Gary Marcus - Marcus on AIgarymarcus.substack.com
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Explore | alphaXivalphaxiv.org
- As Rocks May Think | Eric Jangevjang.com
- Announcing ReasoningLens — Visualizing and Diagnosing LLM Reasoning at a Glancehuggingface.co
- Generative AI's Act o1: The Reasoning Era Begins | Sequoia Capitalsequoiacap.com