Why We Think | Lil'Log
Special thanks to John Schulman for a lot of super valuable feedback and direct edits on this post. Test time compute (Graves et al. 2016, Ling, et al. 2017, Cobbe et al. 2021) and Chain-of-thought (CoT) (Wei et al. 2022, Nye et al. 2021), have led to significant improvements in model performance, while raising many research questions. This post aims to review recent developments in how to effectively use test-time compute (i.e. “thinking time”) and why it helps.
Table of Contents Motivation Analogy to Psychology Computation as a Resource Latent Variable Modeling Thinking in Tokens Branching and Editing Parallel Sampling Sequential Revision RL for Better Reasoning External Tool Use Thinking Faithfully Does the Model Tell What it Thinks Faithfully Optimization Pressure on CoT: Good or Bad? Thinking in Continuous Space Recurrent Architecture Thinking Tokens Thinking as Latent Variables Expectation-Maximization Iterative Learning Scaling Laws for Thinking Time What's for Future Citation References Special thanks to John Schulman for a lot of super valuabl
Explore this link on the map →saved by
- Elizabeth Qiu
- Tasha Pais
- Gabby Chan
- Aryan Naik
- Jennifer Zhao
- Sarah Pan
- Aaron Pham
- Laerdon Kim
- Rikard Saqe
- surya
- Sahil Jain
- Yudhister Joel Kumar
related reading
- DeepSeek-R1arxiv.org
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- [2510.24941] Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thoughtarxiv.org
- Explore | alphaXivalphaxiv.org
- o1 and Reasoning | AndoLogsblog.ando.ai
- Optimizing LLM Test-Time Compute Involves Solving a Meta-RL Problem – Machine Learning Blog | ML@CMU | Carnegie Mellon Universityblog.ml.cmu.edu
- Reasoning as Trajectoriesslhleosun.github.io
- Explore | alphaXivalphaxiv.org
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Vestigial reasoning in RL — LessWronglesswrong.com