Training VLM for CUA - Tzafon
How Tzafon trains Vision Language Models for Computer Use Agents using reinforcement learning, addressing the limitations of SFT and improving generalization across environments.
Research 2026-02-26 Training VLM for CUA How Tzafon trains Vision Language Models for Computer Use Agents using reinforcement learning, addressing the limitations of SFT and improving generalization across environments. Contributors: Nikita Khomich*, Leopold Pluto Hermansson*, David Dinucu Jianu, Ido Hakimi, Yerniyaz Nurgabylov, Noga Bregman, Simon Koser, Noah Löfquist, Mark Rogers *Core contributors Many have tried to solve the task of getting an LLM to use a computer. Until recently, these have all mostly relied on SFT. The reason for this is that it’s easy to define, e.g. just figure out th
saved by
related reading
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io
- Sporks of AGIsergeylevine.substack.com
- 2310.12921.pdfarxiv.org
- Composer2.pdfcursor.com
- RL-VLM-F: Reinforcement Learning from Vision Language Foundation Model Feedbackarxiv.org
- Explore | alphaXivalphaxiv.org
- SimpleVLA-RL: Scaling VLA Training via Reinforcement Learningalphaxiv.org
- To Understand Language is to Understand Generalization | Eric Jangevjang.com
- Language Models can Solve Computer Tasksarxiv.org
- RT-2: Vision-Language-Action Modelsrobotics-transformer2.github.io
- How Claude Performs on Robotics Tasks \ Anthropicanthropic.com
- Emergence of Human to Robot Transfer in Vision-Language-Action Modelspi.website