Anya Singh
1 followers · 3 following · 301 views
on the atlas — 27
- Voyager | An Open-Ended Embodied Agent with Large Language Models4 savers
- Precise Manipulation with Efficient Online RL5 savers
- NVlabs/AutoGaze: AutoGaze automatically removes redundant patches in a video, reducing #tokens in ViT/MLLM by 4x-100x. ·1 savers
- PlayWorld: Learning Robot World Models from Autonomous Play1 savers
- [2603.12254] Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing1 savers
- [2603.08546] Interactive World Simulator for Robot Policy Training and Evaluation1 savers
- Sparse Autoencoders Reveal Interpretable and Steerable Features in VLA Models1 savers
- [2509.22407] EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transfer1 savers
- FASTER1 savers
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels3 savers
- [2510.04371] Speculative Actions: A Lossless Framework for Faster Agentic Systems1 savers
- [2410.00079] Interactive Speculative Planning: Enhance Agent Efficiency through Co-design of System and User Interface1 savers
- [2603.05438] Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model1 savers
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning2 savers
- WebArena-Infinity: Generating Browser Environments with Verifiable Tasks at Scale1 savers
- TRELLIS.2: Native and Compact Structured Latents for 3D Generation2 savers
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AI4 savers
- Six Things I Learned Watching a Robotics Startup Die from the Inside | Rui Xu14 savers
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdf1 savers
- 45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdf1 savers
- Akash Bajwa on X: "RL Environments with Scale AI" / X1 savers
- [2306.00937] STEVE-1: A Generative Model for Text-to-Behavior in Minecraft1 savers
- [2602.10556] LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transfer1 savers
- [2206.11795] Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos2 savers
- Interviewing tips20 savers
- The First Fully General Computer Action Model | blog35 savers
- Power in the Age of Intelligence3 savers
highlights — 4
To solve this, each video prediction is conditioned on the action being executed currently, ensuring a continuous trajectory.
Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIModel inference takes time, but the physical world does not wait for the model to decide. Therefore, we overlap inference and action execution to ensure continuous control, as depicted in Figure 3. Each video prediction is long enough to cover the next prediction's latency.
Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AITranslating Video to Action with Inverse Dynamics Models The second main component of our system is an inverse dynamics model, which performs video-to-action translation: given a predicted video, it produces the precise robot motor signals needed to re-enact the depicted actions. Causal action prediction — as in a typical robot policy — predicts future actions conditioned on the past and thus requires modeling behavior and decision-making. Behavior may be arbitrarily complex, and for generalist behavior, may require a data scale infeasible to collect on robots. In contrast, non-causal video-to…
Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIAll of that data suggests increasing concentration. Through this lens, the SaaSpocalypse (the violent sell-off in software stocks) is less about software writ large dying, and more about point solution software finally facing economic gravity. They are no longer getting a free pass simply for having a good business model.
Power in the Age of Intelligence