feat: algorithm abstraction — named algorithm classes + inline frozen-model references (grpo, opd, sft_distill, self_distill, echo) by hallerite · Pull Request #2746 · PrimeIntellect-ai/prime-rl
NoteRebased onto verifiers-v1. This PR now targets main after the vf-v1 ⇆ nano bridge (#2742) merged. The branch is a clean 2-commit delta on main (the ~80-file algorithm-abstraction change), repla...
feat: algorithm abstraction — named algorithm classes + inline frozen-model references (grpo, opd, sft_distill, self_distill, echo) by hallerite · Pull Request #2746 · PrimeIntellect-ai/prime-rl · GitHub Skip to content You signed in with another tab or window. Reload to refresh your session. You signed out in another tab or window. Reload to refresh your session. You switched accounts on another tab or window. Reload to refresh your session. Dismiss alert {{ message }} Uh oh! There was an error while loading. Please reload this page . PrimeIntellect-ai / prime-rl Public Notifications You must
saved by
related reading
- 1b44b878bb782e6954cd888628510e90-Paper-Conference.pdfproceedings.neurips.cc
- RL at 1T Scale: prime-rl Performance Deep Diveprimeintellect.ai
- Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsnoahziems.com
- Composer2.pdfcursor.com
- Interactive Visualization of RL Algorithms for LLM Trainingzcy233035.github.io
- Algorithms — Ray 2.55.1docs.ray.io
- The 37 Implementation Details of Proximal Policy Optimization · The ICLR Blog Trackiclr-blog-track.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- GRPO++: Tricks for Making RL Actually Workcameronrwolfe.substack.com
- Keep the Tokens Flowing: Lessons from 16 Open-Source RL Librarieshuggingface.co
- Tinkerthinkingmachines.ai
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai