Static HTML snapshot of the model record for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Model | I-JEPA |
| Lab / Organization | Meta FAIR |
| Category | Self-Supervised World Model |
| Subtype | Image Joint Embedding Predictive Architecture |
| World Model Type | Self-supervised visual world model |
| Primary Domain | Image understanding |
| Architecture | Vision Transformer with asymmetric masking and prediction in representation space |
| Modality | Image |
| Training Method | Self-supervised prediction of masked image regions in abstract representation space |
| Status | active |
| Year | 2023 |
| Performance Index | 61/100 (medium confidence, v1.1) |
Main editorial body preserved directly in static HTML.
I-JEPA is the image-domain instantiation of Yann LeCun's JEPA framework. Instead of reconstructing pixels or using contrastive objectives, I-JEPA predicts the abstract representation of a target image block from the representation of a context block. This approach learns semantic visual features that capture high-level structure without being distracted by pixel-level details, following LeCun's vision for non-generative self-supervised world models.
I-JEPA is a self-supervised visual world model developed by Meta in 2023 for image understanding.
Short extractable facts for answer engines and no-JS readers.
| Signal | Value |
|---|---|
| Definition | I-JEPA is a self-supervised visual world model developed by Meta in 2023 for image understanding. |
| Short Description | Image Joint Embedding Predictive Architecture: learns visual representations by predicting abstract image regions without pixel reconstruction. |
| Benchmark Rows | 1 |
| FAQ Entries | 1 |
| Related Models | 1 |
| Related Guides | 0 |
| Related Research Topics | 0 |
| Last Updated | 2026-03-10 |
Key capabilities associated with this model.
Representative applications attached to this model record.
Balanced assessment surfaced in static HTML.
Published benchmark evidence attached to this model record.
| Benchmark | Metric | Result | Source |
|---|---|---|---|
| ImageNet Linear Probe | Top-1 Accuracy 81.1 % | Competitive with MAE | Source |
Primary references preserved in static HTML for citation extraction.
| Reference | Link |
|---|---|
| Assran et al., 2023. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture. CVPR 2023. | Open source |
Nearby models linked from the current editorial record.
| Model | Category | World Model Type | Index v1.1 |
|---|---|---|---|
| V-JEPA | Self-Supervised World Model | Self-supervised visual world model | 70/100 |
Side-by-side comparisons already connected to this model.
| Comparison | Matchup | Summary |
|---|---|---|
| V-JEPA vs Video Generation Models | V-JEPA (Meta) vs Video Generation Models (Sora, Cosmos) | V-JEPA and video generation models like Sora both learn from video, but follow opposite philosophies: V-JEPA predicts in abstract representation space without generating pixels, while video generation models focus on producing realistic pixel outputs. |
| I-JEPA vs MAE (Masked Autoencoders) | I-JEPA (Meta FAIR) vs MAE (Meta / He et al.) | I-JEPA and MAE are both self-supervised image learning methods, but they follow opposite philosophies: I-JEPA predicts in abstract representation space, while MAE reconstructs masked pixels. |
| V-JEPA vs I-JEPA | V-JEPA vs I-JEPA | Both implement Yann LeCun's JEPA framework for self-supervised learning, but V-JEPA operates on video (temporal dynamics) while I-JEPA operates on static images (spatial structure). |
| 3D-VLA vs I-JEPA | 3D-VLA vs I-JEPA | Two approaches to learning representations for embodied intelligence: 3D-VLA combines 3D perception with language-conditioned action planning, while I-JEPA learns abstract visual representations through self-supervised prediction in latent space. |
| V-JEPA vs NVIDIA Cosmos | V-JEPA vs NVIDIA Cosmos | Two foundation-scale approaches to world understanding: V-JEPA learns predictive video representations through self-supervised masking, while Cosmos builds a full-stack world simulation platform for physical AI. |
| LWM vs V-JEPA | Large World Model (LWM) vs V-JEPA | Two approaches to learning world understanding from video. LWM uses autoregressive prediction over million-length sequences, while V-JEPA predicts abstract latent representations without pixel reconstruction. |
| V-JEPA 2 vs V-JEPA | V-JEPA 2 vs V-JEPA | V-JEPA 2 dramatically scales up Meta FAIR's self-supervised video world model, achieving state-of-the-art visual understanding and zero-shot robot control, capabilities V-JEPA didn't demonstrate. |
| V-JEPA 2 vs I-JEPA | V-JEPA 2 vs I-JEPA | Two milestones of the JEPA roadmap: I-JEPA established self-supervised image representation by predicting in latent space; V-JEPA 2 extends the paradigm to video at foundation scale and demonstrates zero-shot robot control. |
FAQ answers rendered directly into static HTML for extractable responses.
I-JEPA operates on static images, predicting masked regions. V-JEPA extends this to video, predicting future frame representations, adding temporal dynamics.
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.
This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-03-10.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary model and lab sources embedded in static HTML.