✳flâneur — a map of the web's best reading
Jeff Brown
0 followers · 95 views
Open this reading profile →
on the atlas — 10
- [2603.21972] Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe1 savers
- The State of Reinforcement Learning for LLM Reasoning1 savers
- A Guide to Reinforcement Learning Post-Training for LLMs: PPO, DPO, GRPO, and Beyond1 savers
- rlhfbook.com/book.pdf1 savers
- Electromagnetic field-inducible in vivo gene switch for remote spatiotemporal control of gene expression: Cell1 savers
- 5D parallelism in LLM training - gdymind's Blog1 savers
- The Ultra-Scale Playbook - a Hugging Face Space by nanotron1 savers
- Curius / Onboarding2507 savers
- The Smol Training Playbook: The Secrets to Building World-Class LLMs - a Hugging Face Space by HuggingFaceTB13 savers
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shi8 savers