o1: A Technical Primer — LessWrong
lesswrong.com · 4,367 words · saved by 1 readers
> TL;DR: In September 2024, OpenAI released o1, its first "reasoning model". This model exhibits remarkable test-time scaling laws, which complete a…
x o1: A Technical Primer — LessWrong Recursive Self-Improvement Scaling Laws Summaries AI Frontpage 175 o1: A Technical Primer by Jesse Hoogland 9th Dec 2024 AI Alignment Forum Linkpost for www.youtube.com 10 min read 19 175 Ω 65 TL;DR : In September 2024, OpenAI released o1, its first "reasoning model". This model exhibits remarkable test-time scaling laws , which complete a missing piece of the Bitter Lesson and open up a new axis for scaling compute. Following Rush and Ritter (2024) and Brown ( 2024a , 2024b ), I explore four hypotheses for how o1 works and discuss some implications for fut
related reading
- o1 and Reasoning | AndoLogsblog.ando.ai
- Reverse engineering OpenAI’s o1interconnects.ai
- Late Takes on OpenAI o1alexirpan.com
- [2412.14135] Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspectivearxiv.org
- As Rocks May Think | Eric Jangevjang.com
- o3 — LessWronglesswrong.com
- Learning to reason with LLMs | OpenAIopenai.com
- OpenAI o1 Results on ARC-AGI-Pub | ARC Prizearcprize.org
- DeepSeek-R1arxiv.org
- [2501.19393] s1: Simple test-time scalingarxiv.org
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- Open-R1: a fully open reproduction of DeepSeek-R1huggingface.co