Recent LLMs can do 2-hop and 3-hop latent (no-CoT) reasoning on natural facts — LessWrong
Prior work has examined 2-hop latent (by "latent" I mean: the model must answer immediately without any Chain-of-Thought) reasoning and found that LLM performance was limited aside from spurious successes (from memorization and shortcuts). An example 2-hop question is: "What element has atomic number (the age at which Tesla died)?". I find that recent LLMs can now do 2-hop and 3-hop latent reasoning with moderate accuracy. I construct a new dataset for evaluating n-hop latent reasoning on natural facts (as in, facts that LLMs already know). On this dataset, I find that Gemini 3 Pro gets 60% of 2-hop questions right and 34% of 3-hop questions right. Opus 4 performs better than Opus 4.5 at this task; Opus 4 gets 31% of 2-hop questions right and 7% of 3-hop questions right. All models I evaluate have chance or near chance accuracy on 4-hop questions. Older models perform much worse; for instance, GPT-4 gets 9.7% of 2-hop questions right and 3.9% of 3-hop questions right. I believe this ne
x Recent LLMs can do 2-hop and 3-hop latent (no-CoT) reasoning on natural facts — LessWrong AI Capabilities Language Models (LLMs) AI Frontpage 2026 Top Fifty: 14 % 129 Recent LLMs can do 2-hop and 3-hop latent (no-CoT) reasoning on natural facts by ryan_greenblatt 1st Jan 2026 AI Alignment Forum 3 min read 11 129 Ω 53 Prior work has examined 2-hop latent (by "latent" I mean: the model must answer immediately without any Chain-of-Thought) reasoning and found that LLM performance was limited aside from spurious successes (from memorization and shortcuts). An example 2-hop question is: "What ele
Explore this link on the map →saved by
related reading
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Reasoning Models Reason Well, Until They Don'tarxiv.org
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Explore | alphaXivalphaxiv.org
- Unlocking the Working Memory of Large Language Models for Latent Reasoningarxiv.org
- Faithful Reasoning (with LLMs)arxiv.org
- Reasoning as Trajectoriesslhleosun.github.io
- 2025.acl-long.896.pdfaclanthology.org
- [2507.06203] A Survey on Latent Reasoningarxiv.org
- Do Large Language Models (LLMs) reason? | Shapedshaped.ai
- [2504.10615] Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Modelsarxiv.org
- Prompt Repetition Improves Non-Reasoning LLMsarxiv.org