Reinforcement learning, AI, and general intelligence
A disclaimer: nothing that I say here is representing any organization other than Artificial Fintelligence. These are my views, and mine alone, although I hope that you share them after reading. Frontier labs are spending, in the aggregate, $100s of millions of dollars annually on data acquisition, leading to a number of startups selling data to them (Mercor, Scale, Surge, etc). The novel data, combined with reinforcement learning (RL) techniques, represents the most clear avenue to improvement, and to AGI. I am firmly convinced that scaling up RL techniques will lead to excellent products, and, eventually, AGI. A primary source of improvement over the last decade has been scale, as the industry has discovered one method after another that allows us to convert money into intelligence. First, bigger models. Then, more data (thereby making Alexandr Wang very rich). And now, RL. Thanks for reading Artificial Fintelligence! Subscribe for free to receive new posts and support my work. RL is
Reinforcement learning and general intelligence Epsilon random is not enough Finbarr Timbers Jun 05, 2025 64 4 3 Share A disclaimer: nothing that I say here is representing any organization other than Artificial Fintelligence. These are my views, and mine alone, although I hope that you share them after reading. Frontier labs are spending, in the aggregate, $100s of millions of dollars annually on data acquisition, leading to a number of startups selling data to them (Mercor, Scale, Surge, etc). The novel data, combined with reinforcement learning (RL) techniques, represents the most clear ave
Explore this link on the map →related reading
- Just Ask for Generalization | Eric Jangevjang.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- The Only Important Technology Is The Internet - Kevin Lukevinlu.ai
- The Era of Experience Paper.pdfstorage.googleapis.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- DeepSeek-R1arxiv.org
- State of RL for reasoning LLMs | A. Weersaweers.de
- GenAI Handbookgenai-handbook.github.io
- Q-learning is not yet scalableseohong.me
- How to scale RL to 10^26 FLOPs - by Jack Morrisblog.jxmo.io
- LLM Resourcesforrestbicker.com
- rlhfbook.com/book.pdfrlhfbook.com