The Annotated JEPA | Elements of a Vector Space
elonlit.com · 9,376 words · saved by 5 readers
An annotated walkthrough of Joint Embedding Predictive Architectures.
--> This post is a step-by-step, annotated, from-scratch walkthrough of Joint Embedding Predictive Architectures, or JEPAs. The goal is to do for JEPA what The Annotated Transformer did for the Transformer: build the full object, explain every moving part, and end with a working training loop. JEPA is Yann LeCun's proposed answer to a fundamental question in self-supervised learning: how do you train a model to understand the world without labels, without collapsing to trivial solutions, and without wasting capacity on irrelevant details? The answer, elegant in principle and subtle in practice
saved by
related reading
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixelsle-wm.github.io
- [2509.14252] LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architecturesarxiv.org
- The first AI model based on Yann LeCun’s vision for more human-like AIai.facebook.com
- V-JEPA: The next step toward advanced machine intelligenceai.meta.com
- the j-lens: finding an llm's unspoken concepts / chiragctxnn.github.io
- arxiv.org/pdf/2511.08544arxiv.org
- Causal-JEPA: Learning World Models through Object-Level Latent Maskingarxiv.org
- [2506.09985] V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planningarxiv.org
- [2603.14482] V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learningarxiv.org
- Verbalizable Representations Form a Global Workspace in Language Modelstransformer-circuits.pub
- [2602.10098] VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Modelarxiv.org
- pdfopenreview.net