✳flâneur — a map of the web's best reading
o1 and Reasoning | AndoLogs
blog.ando.ai · 17,482 words · saved by 1 readers
Organizing research presumably related to OpenAI's o1 reasoning model.
Curious about o1 and “reasoning” in LLMs? Then this post is for you! Here, I attempt to aggregate and organize research related to OpenAI’s o1 . This post is heavily inspired by this awesome-o1 repo/talk [ 1 ] and O1-Journey [ 2 ] for sources and insights grounded in the community’s current consensus, but I’ve also added my own thoughts and interpretations. Disclaimer: All of the thoughts presented here are speculative based on only the publicly available information about o1. I will continue to update this post as more information becomes available, but none of it will come from official or c
Explore this link on the map →saved by
related reading
- VRPRM: Process Reward Modeling via Visual Reasoningarxiv.org
- o1: A Technical Primer — LessWronglesswrong.com
- DeepSeek-R1arxiv.org
- Learning to reason with LLMs | OpenAIopenai.com
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- DeepSeek R1's recipe to replicate o1 and the future of reasoning LMsinterconnects.ai
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Explore | alphaXivalphaxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- Why reasoning models will generalize - by Nathan Lambertinterconnects.ai
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Can activation verbalizers surface an internal chain of thought? — LessWronglesswrong.com