✳flâneur — a map of the web's best reading
The Second Half – Shunyu Yao – 姚顺雨
ysymyth.github.io · 2,432 words · saved by 13 readers
tldr: We’re at AI’s halftime.
The Second Half tldr: We’re at AI’s halftime. For decades, AI has largely been about developing new training methods and models. And it worked: from beating world champions at chess and Go, surpassing most humans on the SAT and bar exams, to earning IMO and IOI gold medals. Behind these milestones in the history book — DeepBlue, AlphaGo, GPT-4, and the o-series — are fundamental innovations in AI methods: search, deep RL, scaling, and reasoning. Things just get better over time. So what’s suddenly different now? In three words: RL finally works. More precisely: RL finally generalizes. After se
Explore this link on the map →saved by
- Aryan Naik
- Hangyul Lyna Kim
- Lydia Nottingham
- surya
- Jirat C
- Dex Ter
- Enrique Moran
- Yudhister Joel Kumar
- Eric Huang
- Nathan Chen
- Sarhaan Gulati
- Rohan Kanti
related reading
- Composer2.pdfcursor.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- DeepSeek-R1arxiv.org
- Clarifying and predicting AGI — LessWronglesswrong.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- What I've Learned About AI in the Past Two Months.sheracaolity.ghost.io
- As Rocks May Think | Eric Jangevjang.com
- [2502.19402] General Reasoning Requires Learning to Reason from the Get-goar5iv.labs.arxiv.org
- Thoughts on the impact of RLHF research — LessWronglesswrong.com
- Are AI benchmarks doomed? - by Anson Ho and Greg Burnhamepochai.substack.com
- How to scale RL to 10^26 FLOPs - by Jack Morrisblog.jxmo.io