✳flâneur — a map of the web's best reading
o1: A Technical Primer — LessWrong
lesswrong.com · 4,367 words · saved by 1 readers
> TL;DR: In September 2024, OpenAI released o1, its first "reasoning model". This model exhibits remarkable test-time scaling laws, which complete a…
x o1: A Technical Primer — LessWrong Recursive Self-Improvement Scaling Laws Summaries AI Frontpage 175 o1: A Technical Primer by Jesse Hoogland 9th Dec 2024 AI Alignment Forum Linkpost for www.youtube.com 10 min read 19 175 Ω 65 TL;DR : In September 2024, OpenAI released o1, its first "reasoning model". This model exhibits remarkable test-time scaling laws , which complete a missing piece of the Bitter Lesson and open up a new axis for scaling compute. Following Rush and Ritter (2024) and Brown ( 2024a , 2024b ), I explore four hypotheses for how o1 works and discuss some implications for fut
Explore this link on the map →related reading
- o1 and Reasoning | AndoLogsblog.ando.ai
- Late Takes on OpenAI o1alexirpan.com
- o3 — LessWronglesswrong.com
- Learning to reason with LLMs | OpenAIopenai.com
- OpenAI o1 Results on ARC-AGI-Pub | ARC Prizearcprize.org
- DeepSeek-R1arxiv.org
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- What I've Learned About AI in the Past Two Months.sheracaolity.ghost.io
- Did OpenAI Just Solve Abstract Reasoning?aiguide.substack.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- How to scale RL to 10^26 FLOPs - by Jack Morrisblog.jxmo.io
- Why reasoning models will generalize - by Nathan Lambertinterconnects.ai