The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity - Apple Machine Learning Research
Authors Parshin Shojaee*†, Iman Mirzadeh*, Keivan Alizadeh, Maxwell Horton, Samy Bengio, Mehrdad Farajtabar View publication Copy Bibtex Recent generations of frontier language models have introduced Large Reasoning Models (LRMs) that generate detailed thinking processes before providing answers. While these models demonstrate improved performance on reasoning benchmarks, their fundamental capabilities, scaling properties, and limitations remain insufficiently understood. Current evaluations primarily focus on established mathematical and coding benchmarks, emphasizing final answer accuracy. However, this evaluation paradigm often suffers from data contamination and does not provide insights into the reasoning traces’ structure and quality. In this work, we systematically investigate these gaps with the help of controllable puzzle environments that allow precise manipulation of compositional complexity while maintaining consistent logical structures. This setup enables the analysis of
AuthorsParshin Shojaee*†, Iman Mirzadeh*, Keivan Alizadeh, Maxwell Horton, Samy Bengio, Mehrdad Farajtabar Recent generations of frontier language models have introduced Large Reasoning Models (LRMs) that generate detailed thinking processes before providing answers. While these models demonstrate improved performance on reasoning benchmarks, their fundamental capabilities, scaling properties, and limitations remain insufficiently understood. Current evaluations primarily focus on established mathematical and coding benchmarks, emphasizing final answer accuracy. However, this evaluation…
saved by
related reading
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Reasoning Models Reason Well, Until They Don'tarxiv.org
- As Rocks May Think | Eric Jangevjang.com
- DeepSeek-R1arxiv.org
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazinequantamagazine.org
- Explore | alphaXivalphaxiv.org
- Training Large Language Models to Reason in a Continuous Latent Spacearxiv.org
- Announcing ReasoningLens — Visualizing and Diagnosing LLM Reasoning at a Glancehuggingface.co
- [2406.02061] Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Modelsarxiv.org
- 'The illusion of thinking': Apple research finds AI models collapse and give up with hard puzzles | Mashablemashable.com
- [2603.07267] How to Steal Reasoning Without Reasoning Tracesarxiv.org