flâneur — a map of the web's best reading

Recent LLMs can do 2-hop and 3-hop latent (no-CoT) reasoning on natural facts — LessWrong

lesswrong.com · 5,861 words · saved by 1 readers

Prior work has examined 2-hop latent (by "latent" I mean: the model must answer immediately without any Chain-of-Thought) reasoning and found that LLM performance was limited aside from spurious successes (from memorization and shortcuts). An example 2-hop question is: "What element has atomic number (the age at which Tesla died)?". I find that recent LLMs can now do 2-hop and 3-hop latent reasoning with moderate accuracy. I construct a new dataset for evaluating n-hop latent reasoning on natural facts (as in, facts that LLMs already know). On this dataset, I find that Gemini 3 Pro gets 60% of 2-hop questions right and 34% of 3-hop questions right. Opus 4 performs better than Opus 4.5 at this task; Opus 4 gets 31% of 2-hop questions right and 7% of 3-hop questions right. All models I evaluate have chance or near chance accuracy on 4-hop questions. Older models perform much worse; for instance, GPT-4 gets 9.7% of 2-hop questions right and 3.9% of 3-hop questions right. I believe this ne

x Recent LLMs can do 2-hop and 3-hop latent (no-CoT) reasoning on natural facts — LessWrong AI Capabilities Language Models (LLMs) AI Frontpage 2026 Top Fifty: 14 % 129 Recent LLMs can do 2-hop and 3-hop latent (no-CoT) reasoning on natural facts by ryan_greenblatt 1st Jan 2026 AI Alignment Forum 3 min read 11 129 Ω 53 Prior work has examined 2-hop latent (by "latent" I mean: the model must answer immediately without any Chain-of-Thought) reasoning and found that LLM performance was limited aside from spurious successes (from memorization and shortcuts). An example 2-hop question is: "What ele

Explore this link on the map →

saved by

related reading