45d74e190008c7bff2845ffc8e3facd3-Paper-Conference.pdf
proceedings.iclr.cc · 13,038 words · saved by 1 readers
N/A
# link_1fup4l9vx3f.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=true - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20250408014223Z - Creator=LaTeX with hyperref - ModDate=D:20250408014223Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.26 (TeX Live 2024) kpathsea version 6.4.0 - Producer=pdfTeX-1.40.26 - Trapped=False ## Contents ### Page 1 Published as a conference paper at ICLR 2025LATENT ACTION PRETRAINING FROM VIDEOSSeonghyeon Ye1∗† Joel Jang2∗‡Byeongguk Jeon1 Sejune Joo1 Jianwei Ya
saved by
related reading
- [2602.10556] LAP: Language-Action Pre-Training Enables Zero-shot Cross-Embodiment Transferarxiv.org
- e5b5c402bb7bd5e60bede6961d6fe39e-Paper-Conference.pdfproceedings.iclr.cc
- [2509.22407] EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transferarxiv.org
- [2603.08546] Interactive World Simulator for Robot Policy Training and Evaluationarxiv.org
- [2206.11795] Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videosarxiv.org
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org
- Causal Video Models Are Data-Efficient Robot Policy Learners | Rhoda AIrhoda.ai
- [2410.11758] Latent Action Pretraining from Videosarxiv.org
- The First Fully General Computer Action Model | blogsi.inc
- [2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Modelarxiv.org
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- State of Vision-Language-Action (VLA) Research at ICLR 2026 – Moritz Reussmbreuss.github.io