V-JEPA: The next step toward advanced machine intelligence
As humans, much of what we learn about the world around us—particularly in our early stages of life—is gleaned through observation. Take Newton’s third law of motion: Even an infant (or a cat) can intuit, after knocking several items off a table and observing the results, that what goes up must come down. You don’t need hours of instruction or to read thousands of books to arrive at that result. Your internal world model—a contextual understanding based on a mental model of the world—predicts these consequences for you, and it’s highly efficient. “V-JEPA is a step toward a more grounded understanding of the world so machines can achieve more generalized reasoning and planning,” says Meta’s VP & Chief AI Scientist Yann LeCun, who proposed the original Joint Embedding Predictive Architectures (JEPA) in 2022. “Our goal is to build advanced machine intelligence that can learn more like humans do, forming internal models of the world around them to learn, adapt, and forge plans efficiently
V-JEPA: The next step toward advanced machine intelligence Products AI Research Resources About Get Llama Try Meta AI Research V-JEPA: The next step toward Yann LeCun’s vision of advanced machine intelligence (AMI) February 15, 2024 • 3 minute read Takeaways: Today, we’re publicly releasing the Video Joint Embedding Predictive Architecture (V-JEPA) model, a crucial step in advancing machine intelligence with a more grounded understanding of the world. This early example of a physical world model excels at detecting and understanding highly detailed interactions between objects. In the spirit o
Explore this link on the map →saved by
related reading
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org
- The Annotated JEPA | Elements of a Vector Spaceelonlit.com
- The first AI model based on Yann LeCun’s vision for more human-like AIai.facebook.com
- how we accidentally solved robotics by watching 1 million hours of YouTube – atharva's blogksagar.bearblog.dev
- [2603.14482] V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learningarxiv.org
- The First Fully General Computer Action Model | blogsi.inc
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixelsle-wm.github.io
- Video models are zero-shot learners and reasonersarxiv.org
- The flavor of the bitter lesson for computer vision - Vincent Sitzmannvincentsitzmann.com
- The Model That Dreams the Worldmoe-capital.com
- Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blogdeveloper.nvidia.com
- World Models | Rohit Bandarurohitbandaru.github.io