The Second Half – Shunyu Yao – 姚顺雨
ysymyth.github.io · 2,432 words · saved by 17 readers
tldr: We’re at AI’s halftime.
The Second Half tldr: We’re at AI’s halftime. For decades, AI has largely been about developing new training methods and models. And it worked: from beating world champions at chess and Go, surpassing most humans on the SAT and bar exams, to earning IMO and IOI gold medals. Behind these milestones in the history book — DeepBlue, AlphaGo, GPT-4, and the o-series — are fundamental innovations in AI methods: search, deep RL, scaling, and reasoning. Things just get better over time. So what’s suddenly different now? In three words: RL finally works. More precisely: RL finally generalizes. After se
saved by
- Aryan Naik
- Hangyul Lyna Kim
- Lydia Nottingham
- surya
- Jirat C
- Dex Ter
- Enrique Moran
- Yudhister Joel Kumar
- Eric Huang
- Nathan Chen
- Sarhaan Gulati
- Rohan Kanti
related reading
- As Rocks May Think | Eric Jangevjang.com
- Composer2.pdfcursor.com
- Ilya Sutskever — We're moving from the age of scaling to the age of researchdwarkesh.com
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- AI in 2025: gestalt — LessWronglesswrong.com
- The Extreme Inefficiency of RL for Frontier Models - Toby Ordtobyord.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Clarifying and predicting AGI — LessWronglesswrong.com
- Things I learned at OpenAI - by Karina Nguyen - sémaphoresemaphore.substack.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- Questions about the Future of AI - by Dwarkesh Pateldwarkesh.com