Reverse engineering OpenAI’s o1 - by Nathan Lambert
On Episode 33 of The Retort, we discussed the Reflection 70b drama and more motivations for model specs and transparency. OpenAI released their new reasoning system, o1, building on the early successes of Q* and more recently the rumors of Strawberry, to ship a new mode of interacting with AI on challenging tasks. o1 is a system designed by training new models on long reasoning chains, with lots of reinforcement learning 🍒, and deploying them at scale. Unlike traditional autoregressive language models, it is doing an online search for the user. It is spending more on inference, which confirms the existence of new scaling laws — inference scaling laws. The release is far from a coherent product. o1 is a prototype. o1 is a vague name (”o for OpenAI”). o1 does not have the clarity of product-market fit that ChatGPT did. o1 is extremely powerful. o1 is different. o1 is a preview of the future of AI. AI’s march forward of progress technically has continued. Its capability overhang has deep
On Episode 33 of The Retort, we discussed the Reflection 70b drama and more motivations for model specs and transparency. OpenAI released their new reasoning system, o1, building on the early successes of Q* and more recently the rumors of Strawberry, to ship a new mode of interacting with AI on challenging tasks. o1 is a system designed by training new models on long reasoning chains, with lots of reinforcement learning 🍒, and deploying them at scale. Unlike traditional autoregressive language models, it is doing an online search for the user. It is spending more on inference, which…
saved by
related reading
- Learning to reason with LLMs | OpenAIopenai.com
- o1: A Technical Primer — LessWronglesswrong.com
- o1 and Reasoning | AndoLogsblog.ando.ai
- Late Takes on OpenAI o1alexirpan.com
- o1 System Cardassets.ctfassets.net
- [2412.14135] Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspectivearxiv.org
- As Rocks May Think | Eric Jangevjang.com
- OpenAI o1 Results on ARC-AGI-Pub | ARC Prizearcprize.org
- Open-R1: a fully open reproduction of DeepSeek-R1huggingface.co
- o3, Oh Mythezvi.substack.com
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- DeepSeek-R1arxiv.org