Infini-AI-Lab on X: "Weโre excited to release ๐๐ฌ๐ญ๐ซ๐๐ ๐ฅ๐จ๐ฐ, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. ๐ Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: โก ๐.๐ร ๐๐๐ฌ๐ญ๐๐ซ ๐ฆ๐ฎ๐ฅ๐ญ๐ข-๐ฉ๐จ๐ฅ๐ข๐๐ฒ https://t.co/JVthM8iHur" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Home Explore Notifications Chat Grok Premium Money Bookmarks Creator Studio Articles Profile More Post christina @luoluo Post See new posts Conversation Infini-AI-Lab @InfiniAILab Weโre excited to release ๐๐ฌ๐ญ๐ซ๐๐ ๐ฅ๐จ๐ฐ, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: ๐.๐ร ๐๐๐ฌ๐ญ๐๐ซ ๐ฆ๐ฎ๐ฅ๐ญ๐ข-๐ฉ๐จ๐ฅ๐ข๐๐ฒ ๐๐ ๐๐ง๐ญ๐ฌ ๐๐จ๐ฅ๐ฅ๐๐๐จ๐ซ๐๐ญ๐ข๐ฏ๐ ๐๐ ๐ญ๐ซ๐๐ข๐ง๐ข๐ง๐ Achieves comparable or better accuracy than verl-based baseline. ๐๐๐ซ๐จ-๐๐จ๐๐ ๐ฌ๐ฒ๐ฌ๐ญ๐๐ฆ ๐๐ฅ๐๐ฑ๐ข๐๐ข๐ฅ๐ข๐ญ๐ฒ Supports elastic multi-policy training and cross-region rollout across heterogeneous GPUs. โค๐.๐% ๐ฌ๐ฉ๐๐ซ๐ฌ๐ ๐ญ๐ซ๐๐ง๐ฌ๐๐๐ซ ๐๐จ๐ซ ๐ซ๐๐ฆ๐จ๐ญ๐ ๐ซ๐จ๐ฅ๐ฅ๐จ๐ฎ๐ญ Same to @FireworksAI_HQ โs sparse RL transfer design, AstraFlow cuts sync from ~28 GB to ~1.5 GB, with deltas โค1.1% of weights, m
@InfiniAILab: Weโre excited to release ๐๐ฌ๐ญ๐ซ๐๐ ๐ฅ๐จ๐ฐ, an open-source, dataflow-oriented RL system for training multi-agentic and multi-policy LLMs. Built for scalable, flexible, and efficient agent RL, AstraFlow natively enables: ๐.๐ร ๐๐๐ฌ๐ญ๐๐ซ ๐ฆ๐ฎ๐ฅ๐ญ๐ข-๐ฉ๐จ๐ฅ๐ข๐๐ฒ ๐๐ ๐๐ง๐ญ๐ฌ ๐๐จ๐ฅ๐ฅ๐๐๐จ๐ซ๐๐ญ๐ข๐ฏ๐ ๐๐ ๐ญ๐ซ๐๐ข๐ง๐ข๐ง๐ Achieves comparable or better accuracy than verl-based baseline. ๐๐๐ซ๐จ-๐๐จ๐๐ ๐ฌ๐ฒ๐ฌ๐ญ๐๐ฆ ๐๐ฅ๐๐ฑ๐ข๐๐ข๐ฅ๐ข๐ญ๐ฒ Supports elastic multi-policy training and cross-region rollout across heterogeneous GPUs. โค๐.๐% ๐ฌ๐ฉ๐๐ซ๐ฌ๐ ๐ญ๐ซ๐๐ง๐ฌ๐๐๐ซโฆ
Explore this link on the map โsaved by
related reading
- Sonya Huang ๐ฅ on X: "Today's Training Data episode takes us BTS on the infrastructure challenges required to do large RL runs at scale, featuring @ellev3n11 (Composer Lead at @cursor_ai) and @dzhulgakov (Co-Founder at @FireworksAI_HQ). The Cursor team trained Composer 2 on Fireworks by starting with https://t.co/6LLlJlyl8Q" / Xx.com
- Is Frontier Asynchronous RL Solved? โ Luke J. Huangluk-huang.github.io
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Rethinking RL Infra for Agents | B'Logbillxbf.github.io
- Composer2.pdfcursor.com
- State of RL for reasoning LLMs | A. Weersaweers.de
- Forge: Scalable Agent RL Framework and Algorithm - MiniMax News | MiniMaxminimax.io
- Explore | alphaXivalphaxiv.org
- Building Effective AI Agents \ Anthropicanthropic.com
- Building Effective AI Agents \ Anthropicanthropic.com
- PostTrainBenchposttrainbench.com
- [2603.21972] Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipearxiv.org