feat: algorithm abstraction — named algorithm classes + inline frozen-model references (grpo, opd, sft_distill, self_distill, echo) by hallerite · Pull Request #2746 · PrimeIntellect-ai/prime-rl
NoteRebased onto verifiers-v1. This PR now targets main after the vf-v1 ⇆ nano bridge (#2742) merged. The branch is a clean 2-commit delta on main (the ~80-file algorithm-abstraction change), repla...
feat: algorithm abstraction — named algorithm classes + inline frozen-model references (grpo, opd, sft_distill, self_distill, echo) by hallerite · Pull Request #2746 · PrimeIntellect-ai/prime-rl · GitHub Skip to content You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} Uh oh! There was an error while loading. Please reload this page . PrimeIntellect-ai / prime-rl Public Notifications You must
Explore this link on the map →related reading
- Composer2.pdfcursor.com
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- Algorithms — Ray 2.55.1docs.ray.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- A vision researcher’s guide to some RL stuff: PPO & GRPO - Yuge (Jimmy) Shiyugeten.github.io
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com
- SFT, RL, and On-Policy Distillation Through a Distributional Lens | whnrehiew.github.io
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- o1 and Reasoning | AndoLogsblog.ando.ai
- GitHub - labmlai/annotated_deep_learning_paper_implementations: 🧑🏫 60+ Implementations/tutorials of deep learning papers with side-by-side notes 📝; including transformers (original, xl, switch, feedback, vit, ...), optimizers (adam, adagithub.com
- RL Scaling Laws for LLMs - by Cameron R. Wolfe, Ph.D.cameronrwolfe.substack.com
- AlgZoo: uninterpreted models with fewer than 1,500 parameters — LessWronglesswrong.com