Reverse engineering OpenAI’s o1 - by Nathan Lambert
On Episode 33 of The Retort, we discussed the Reflection 70b drama and more motivations for model specs and transparency. OpenAI released their new reasoning system, o1, building on the early successes of Q* and more recently the rumors of Strawberry, to ship a new mode of interacting with AI on challenging tasks. o1 is a system designed by training new models on long reasoning chains, with lots of reinforcement learning 🍒, and deploying them at scale. Unlike traditional autoregressive language models, it is doing an online search for the user. It is spending more on inference, which confirms the existence of new scaling laws — inference scaling laws. The release is far from a coherent product. o1 is a prototype. o1 is a vague name (”o for OpenAI”). o1 does not have the clarity of product-market fit that ChatGPT did. o1 is extremely powerful. o1 is different. o1 is a preview of the future of AI. AI’s march forward of progress technically has continued. Its capability overhang has deep
Explore this link on the map →