[2504.10615] Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models
Abstract:Large language models (LLMs) can perform reasoning computations both internally within their latent space and externally by generating explicit token sequences like chains of thought. Significant progress in enhancing reasoning abilities has been made by scaling test-time compute. However, understanding and quantifying model-internal reasoning abilities - the inferential "leaps" models make between individual token predictions - remains crucial. This study introduces a benchmark (n = 4,000 items) designed to quantify model-internal reasoning in different domains. We achieve this by having LLMs indicate the correct solution to reasoning problems not through descriptive text, but by selecting a specific language of their initial response token that is different from English, the benchmark language. This not only requires models to reason beyond their context window, but also to overrise their default tendency to respond in the same language as the prompt, thereby posing an additional cognitive strain. We evaluate a set of 18 LLMs, showing significant performance variations, with GPT-4.5 achieving the highest accuracy (74.7%), outperforming models like Grok-2 (67.2%), and Llama 3.1 405B (65.6%). Control experiments and difficulty scaling analyses suggest that while LLMs engage in internal reasoning, we cannot rule out heuristic exploitations under certain conditions, marking an area for future investigation. Our experiments demonstrate that LLMs can "think" via latent-space computations, revealing model-internal inference strategies that need further understanding, especially regarding safety-related concerns such as covert planning, goal-seeking, or deception emerging without explicit token traces.
[2504.10615] Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2504.10615 (cs) [Submitted on 14 Apr 2025] Title: Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models Authors: Thilo Hagendorff , Sarah Fabi View a PDF of the paper titled Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large
related reading
- Training Large Language Models to Reason in a Continuous Latent Spacearxiv.org
- [2412.06769] Training Large Language Models to Reason in a Continuous Latent Spacearxiv.org
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- DeepSeek-R1arxiv.org
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- As Rocks May Think | Eric Jangevjang.com
- [2507.06203] A Survey on Latent Reasoningarxiv.org
- Unlocking the Working Memory of Large Language Models for Latent Reasoningarxiv.org
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexitymachinelearning.apple.com
- Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazinequantamagazine.org
- Explore | alphaXivalphaxiv.org