Eric Jang on X: "@rosstaylor90 Great questions. Have been asking myself a lot of the same things. 1. we were "lucky" that early GPT3-4 era models happened to exhibit chain of thought with the right prompt, but on-policy RL makes success a lot more guaranteed. I think the current warning sign in the land" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Home Explore Notifications Chat Grok Premium Bookmarks Creator Studio Articles Profile More Post Dhruv Sheth @dhruvsheth_ Post See new posts Conversation Ross Taylor @rosstaylor90 · 2h Thoughts and held-out answers: - Prior to the RL paradigm, the benchmarks clearly showed reasoning was a problem (eg GPT-4 got 42.5% on MATH). Where are the warning signs right now that show the current recipe is inadequate? Do the right evals even exist yet? - Suppose it is Show more 3 24 2.7K Eric Jang @ericjang11 Great questions. Have been asking myself a lot of the same things. 1. we were "lucky" that early GPT3-4 era models happened to exhibit chain of thought with the right prompt, but on-policy RL makes success a lot more guaranteed. I think the current warning sign in the land of "thinking" is the classic "stock predictor / scientific world model causal confusion problem", where models can overfit knowledge from the f
Eric Jang @ericjang11 Replying to @rosstaylor90 Great questions. Have been asking myself a lot of the same things. 1. we were "lucky" that early GPT3-4 era models happened to exhibit chain of thought with the right prompt, but on-policy RL makes success a lot more guaranteed. I think the current warning sign in the land of "thinking" is the classic "stock predictor / scientific world model causal confusion problem", where models can overfit knowledge from the future in their weights instead of learning to reason from available facts *now* (which is how we plan to use them). they have too much
Explore this link on the map →related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- DeepSeek-R1arxiv.org
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- gpt-4.pdfcdn.openai.com
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Vestigial reasoning in RL — LessWronglesswrong.com
- GRPO is terrible — LessWronglesswrong.com
- Thoughts on the impact of RLHF research — LessWronglesswrong.com
- By Default, GPTs Think In Plain Sight — LessWronglesswrong.com
- o1 and Reasoning | AndoLogsblog.ando.ai
- How Does Claude 4 Think? — Sholto Douglas & Trenton Brickendwarkesh.com