Training Dyna-2 at million-hour scale, repeatably — DYNA
dyna.co · 4,152 words · saved by 1 readers
How we rebuilt every step of our robotics data stack for million-hour scale training.
[ Research ] Training Dyna-2 at million-hour scale, repeatably Category: Research Author: Dyna Robotics Date: August 2026 Read: 22 min Robot or vendorcollectionLanding bucketobject storageTraining-ready dataMCAP episodesTraining manifestcolumnar + mmapGPU clustertraininguploadsingestioncurationloading Figure 1: Flowchart of data lifecycle From collection to training on the GPU cluster, robot data moves through four stages: collection into a landing bucket, ingestion into training-ready episodes, curation into the dataset for a given experiment, and loading into training batches.…
saved by
related reading
- Nikolaus West (@NikolausWest) on Xx.com
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- Multi-Datacenter Training: OpenAI's Ambitious Plan To Beat Google's Infrastructuresemianalysis.com
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- Many Small Steps for Robots, One Giant Leap for Mankindnotboring.co
- State of Robot Learning, December 2025vedder.io
- Training great LLMs entirely from ground up in the wilderness as a startup - Yi Tayyitay.net
- Macrodata Labs — better data for better robotsmacrodata.co
- ACT-1: A Robot Foundation Model Trained on Zero Robot Data | Sunday Robotics | The helpful robotics companysunday.ai
- Reinforcement learning is an infrastructure problemmodal.com
- MAI-Thinking-1: Building a Hill-Climbing Machinemicrosoft.ai
- Keep the Tokens Flowing: Lessons from 16 Open-Source RL Librarieshuggingface.co