The First Fully General Computer Action Model | blog
We trained a model on our 11-million-hour video dataset. Our model can explore complex websites, complete multi-action CAD modeling sequences, and drive a car in the real world, all at 30 FPS.
We designed FDM-1, a foundation model for computer use. FDM-1 is trained on videos from a portion of our 11-million-hour screen recording dataset, which we labeled using an inverse dynamics model that we trained. Our video encoder can compress almost 2 hours of 30 FPS video in only 1M tokens. FDM-1 is the first model with the long-context training needed to become a coworker for CAD, finance, engineering, and eventually ML research, and it consistently improves with scale. It trains and infers directly on video instead of screenshots and can learn unsupervised from the entirety of the internet
Explore this link on the map →saved by
- Pranav
- Justin Wang
- Elizabeth Qiu
- Rajan Agarwal
- Tasha Pais
- Ratan Kaliani
- HudZah
- Emma Guo
- Eshaan Moorjani
- Rishi Kothari
- Neel Redkar
- Aaron Pham
related reading
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdfproceedings.iclr.cc
- Generalist - GEN-0 / Embodied Foundation Models That Scale with Physical Interactiongeneralistai.com
- World Models: Computing the Uncomputablenotboring.co
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- Composer2.pdfcursor.com
- World Models | Rohit Bandarurohitbandaru.github.io
- [2206.11795] Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videosarxiv.org
- [2410.11758] Latent Action Pretraining from Videosarxiv.org
- Explore | alphaXivalphaxiv.org
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- [2206.11795] Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videosarxiv.org