o1 and Reasoning | AndoLogs
blog.ando.ai · 17,482 words · saved by 1 readers
Organizing research presumably related to OpenAI's o1 reasoning model.
Curious about o1 and “reasoning” in LLMs? Then this post is for you! Here, I attempt to aggregate and organize research related to OpenAI’s o1 . This post is heavily inspired by this awesome-o1 repo/talk [ 1 ] and O1-Journey [ 2 ] for sources and insights grounded in the community’s current consensus, but I’ve also added my own thoughts and interpretations. Disclaimer: All of the thoughts presented here are speculative based on only the publicly available information about o1. I will continue to update this post as more information becomes available, but none of it will come from official or c
saved by
related reading
- o1: A Technical Primer — LessWronglesswrong.com
- As Rocks May Think | Eric Jangevjang.com
- DeepSeek-R1arxiv.org
- Learning to reason with LLMs | OpenAIopenai.com
- Reverse engineering OpenAI’s o1interconnects.ai
- [2412.14135] Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspectivearxiv.org
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- Open-R1: a fully open reproduction of DeepSeek-R1huggingface.co
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Explore | alphaXivalphaxiv.org
- Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazinequantamagazine.org