Infini-AI-Lab on X: "Weโre excited to release ๐๐ฌ๐ญ๐ซ๐๐ ๐ฅ๐จ๐ฐ, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. ๐ Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: โก ๐.๐ร ๐๐๐ฌ๐ญ๐๐ซ ๐ฆ๐ฎ๐ฅ๐ญ๐ข-๐ฉ๐จ๐ฅ๐ข๐๐ฒ https://t.co/JVthM8iHur" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Home Explore Notifications Chat Grok Premium Money Bookmarks Creator Studio Articles Profile More Post christina @luoluo Post See new posts Conversation Infini-AI-Lab @InfiniAILab Weโre excited to release ๐๐ฌ๐ญ๐ซ๐๐ ๐ฅ๐จ๐ฐ, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: ๐.๐ร ๐๐๐ฌ๐ญ๐๐ซ ๐ฆ๐ฎ๐ฅ๐ญ๐ข-๐ฉ๐จ๐ฅ๐ข๐๐ฒ ๐๐ ๐๐ง๐ญ๐ฌ ๐๐จ๐ฅ๐ฅ๐๐๐จ๐ซ๐๐ญ๐ข๐ฏ๐ ๐๐ ๐ญ๐ซ๐๐ข๐ง๐ข๐ง๐ Achieves comparable or better accuracy than verl-based baseline. ๐๐๐ซ๐จ-๐๐จ๐๐ ๐ฌ๐ฒ๐ฌ๐ญ๐๐ฆ ๐๐ฅ๐๐ฑ๐ข๐๐ข๐ฅ๐ข๐ญ๐ฒ Supports elastic multi-policy training and cross-region rollout across heterogeneous GPUs. โค๐.๐% ๐ฌ๐ฉ๐๐ซ๐ฌ๐ ๐ญ๐ซ๐๐ง๐ฌ๐๐๐ซ ๐๐จ๐ซ ๐ซ๐๐ฆ๐จ๐ญ๐ ๐ซ๐จ๐ฅ๐ฅ๐จ๐ฎ๐ญ Same to @FireworksAI_HQ โs sparse RL transfer design, AstraFlow cuts sync from ~28 GB to ~1.5 GB, with deltas โค1.1% of weights, m
@InfiniAILab: Weโre excited to release ๐๐ฌ๐ญ๐ซ๐๐ ๐ฅ๐จ๐ฐ, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: ๐.๐ร ๐๐๐ฌ๐ญ๐๐ซ ๐ฆ๐ฎ๐ฅ๐ญ๐ข-๐ฉ๐จ๐ฅ๐ข๐๐ฒ ๐๐ ๐๐ง๐ญ๐ฌ ๐๐จ๐ฅ๐ฅ๐๐๐จ๐ซ๐๐ญ๐ข๐ฏ๐ ๐๐ ๐ญ๐ซ๐๐ข๐ง๐ข๐ง๐ Achieves comparable or better accuracy than verl-based baseline. ๐๐๐ซ๐จ-๐๐จ๐๐ ๐ฌ๐ฒ๐ฌ๐ญ๐๐ฆ ๐๐ฅ๐๐ฑ๐ข๐๐ข๐ฅ๐ข๐ญ๐ฒ Supports elastic multi-policy training and cross-region rollout across heterogeneous GPUs. โค๐.๐% ๐ฌ๐ฉ๐๐ซ๐ฌ๐ ๐ญ๐ซ๐๐ง๐ฌ๐๐๐ซโฆ
saved by
related reading
- RL at 1T Scale: prime-rl Performance Deep Diveprimeintellect.ai
- Is Frontier Asynchronous RL Solved? โ Luke J. Huangluk-huang.github.io
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Keep the Tokens Flowing: Lessons from 16 Open-Source RL Librarieshuggingface.co
- Rethinking RL Infra for Agents | B'Logbillxbf.github.io
- Composer2.pdfcursor.com
- Together AI | The AI Native Cloudtogether.ai
- State of RL for reasoning LLMs | A. Weersaweers.de
- Building Effective AI Agents \ Anthropicanthropic.com
- PostTrainBenchposttrainbench.com
- Forge: Scalable Agent RL Framework and Algorithm - MiniMax News | MiniMaxminimax.io
- [2602.19362] LLMs Can Learn to Reason Via Off-Policy RLarxiv.org