flâneur — a map of the web's best reading

State of RL for reasoning LLMs | A. Weers

aweers.de · 5,052 words · saved by 6 readers

PhD student

State of RL for reasoning LLMs ¶ Reinforcement learning has been one of the most consequential additions to the LLM post-training stack. It was the key ingredient that transformed GPT-3 into InstructGPT [1] , and it has since become central to the current wave of reasoning improvements [2] [3] . The first generation of RL for LLMs was dominated by PPO [4] , a method developed for more conventional RL settings such as Atari games and robotics, but adapted very successfully to RLHF. The second generation, driven by the goal of improving reasoning capabilities, brought another round of algor

Explore this link on the map →

saved by

related reading