2404.15758
arxiv.org · 7,055 words · saved by 1 readers
N/A
Let’s Think Dot by Dot: Hidden Computation in Transformer Language Models Jacob Pfau, William Merrill & Samuel R. Bowman Center for Data Science New York University NY, NY 10012, USA {jp6263,willm,bowman}@nyu.edu arXiv:2404.15758v1 [cs.CL] 24 Apr 2024…
related reading
- [2404.15758] Let's Think Dot by Dot: Hidden Computation in Transformer Language Modelsarxiv.org
- Let's Think Dot by Dot: Hidden Computation in Transformer Language Modelsarxiv.org
- LLMs are (mostly) not helped by filler tokens — LessWronglesswrong.com
- 2402.12875arxiv.org
- A Mathematical Framework for Transformer Circuitstransformer-circuits.pub
- [2402.12875] Chain of Thought Empowers Transformers to Solve Inherently Serial Problemsarxiv.org
- Recent LLMs can use filler tokens or problem repeats to improve (no-CoT) math performanceblog.redwoodresearch.org
- Transformer Circuits Threadtransformer-circuits.pub
- dennyzhou.github.io/LLM-Reasoning-Stanford-CS-25.pdfdennyzhou.github.io
- Thinking like Transformersrush.github.io
- Transformers Provably Learn to Internalize Chain-of-Thoughtarxiv.org
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io