flâneur

Harsh Bhatt on X: "The Math of RL: From Policy Gradients to PPO and GRPO" / X

x.com · 2,525 words · saved by 1 readers

The Math of RL: From Policy Gradients to PPO and GRPO

Reinforcement Learning is becoming very popular since last few years especially in LLMs with the techniques like RLHF (Reinforcement Learning Human Feedback) and RLVR (Reinforcement Learning Verifiable Rewards). RL got the very high attention when DeepMind built and trained AlphaGo, but there came a period when shift moved towards unsupervised technique adversarial loss with GANs, semi-supervised learning with transformers, but as LLMs uplifted the entire deep learning, researchers looked backed into RL. Because of RL we can able to align LLMs using RLHF and make them expert in skills like…

saved by

related reading