Static HTML snapshot of the model record for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Model | V-JEPA 2 |
| Lab / Organization | Meta FAIR |
| Category | Self-Supervised World Model |
| Subtype | Video Prediction Model |
| World Model Type | Joint-embedding predictive world model for video understanding and robot planning |
| Primary Domain | Physical Reasoning & Robotics |
| Architecture | Vision Transformer with joint-embedding predictive architecture, latent space prediction |
| Modality | Video → Latent Predictions |
| Training Method | Self-supervised learning via latent space prediction from video, no pixel reconstruction |
| Status | active |
| Year | 2025 |
| Performance Index | 87/100 (medium confidence, v1.1) |
Main editorial body preserved directly in static HTML.
V-JEPA 2 (Video Joint Embedding Predictive Architecture 2) is a self-supervised foundation world model from Meta FAIR that learns to understand, predict, and plan from video without relying on pixel-level reconstruction. By predicting in a learned latent space rather than pixel space, V-JEPA 2 avoids the computational overhead and irrelevant detail of generative approaches. It achieves state-of-the-art performance on visual understanding benchmarks and demonstrates zero-shot robot control capabilities, planning actions in new environments without task-specific fine-tuning. The model introduces new physical reasoning benchmarks and is fully open-source, marking a major step toward Yann LeCun's vision of world-model-based AI.
V-JEPA 2 is a joint-embedding predictive world model for video understanding and robot planning developed by Meta in 2025 for physical reasoning & robotics.
Short extractable facts for answer engines and no-JS readers.
| Signal | Value |
|---|---|
| Definition | V-JEPA 2 is a joint-embedding predictive world model for video understanding and robot planning developed by Meta in 2025 for physical reasoning & robotics. |
| Short Description | Meta FAIR's self-supervised video world model achieving state-of-the-art visual understanding and enabling zero-shot robot control. |
| Benchmark Rows | 1 |
| FAQ Entries | 2 |
| Related Models | 3 |
| Related Guides | 1 |
| Related Research Topics | 2 |
| Last Updated | 2026-04-10 |
Key capabilities associated with this model.
Representative applications attached to this model record.
Balanced assessment surfaced in static HTML.
Published benchmark evidence attached to this model record.
| Benchmark | Metric | Result | Source |
|---|---|---|---|
| Physical Reasoning | PhyBench Accuracy 89.2 % | State-of-the-art | Source |
Primary references preserved in static HTML for citation extraction.
| Reference | Link |
|---|---|
| Assran et al., 2025. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv:2506.09985. | Open source |
Nearby models linked from the current editorial record.
| Model | Category | World Model Type | Index v1.1 |
|---|---|---|---|
| V-JEPA | Self-Supervised World Model | Self-supervised visual world model | 70/100 |
| I-JEPA | Self-Supervised World Model | Self-supervised visual world model | 61/100 |
| AMI World Model | Foundation World Model | Multimodal generative world model | 38/100 |
Side-by-side comparisons already connected to this model.
| Comparison | Matchup | Summary |
|---|---|---|
| V-JEPA 2 vs V-JEPA | V-JEPA 2 vs V-JEPA | V-JEPA 2 dramatically scales up Meta FAIR's self-supervised video world model, achieving state-of-the-art visual understanding and zero-shot robot control, capabilities V-JEPA didn't demonstrate. |
| V-JEPA 2 vs I-JEPA | V-JEPA 2 vs I-JEPA | Two milestones of the JEPA roadmap: I-JEPA established self-supervised image representation by predicting in latent space; V-JEPA 2 extends the paradigm to video at foundation scale and demonstrates zero-shot robot control. |
| LeWorldModel vs DreamerV3 | LeWorldModel vs DreamerV3 | LeWorldModel revisits LeCun's energy-based JEPA philosophy for control, predicting in latent space without pixel reconstruction. DreamerV3 remains the canonical RSSM-based agent that learns by imagining pixel-grounded rollouts. |
| PlayWorld vs TD-MPC2 | PlayWorld vs TD-MPC2 | Two green-index models for robot decision-making, but with very different operating modes. PlayWorld learns a manipulation-focused world simulator from autonomous play, while TD-MPC2 combines latent dynamics with model-predictive control across a wide multi-task control benchmark suite. |
| Genie 3 vs NVIDIA Cosmos | Genie 3 vs NVIDIA Cosmos | Two green-index frontier systems with different ambitions. Genie 3 is a real-time text-to-world interactive generator, while NVIDIA Cosmos is a broad physical-AI platform optimized for simulation infrastructure, robotics, and industrial world modeling. |
| Genie 3 vs V-JEPA 2 | Genie 3 vs V-JEPA 2 | Two green-index leaders that represent different frontier philosophies. Genie 3 is an interactive generative world model that turns text into playable environments, while V-JEPA 2 is a self-supervised latent predictor optimized for physical reasoning and zero-shot robot planning. |
| NVIDIA Cosmos vs V-JEPA 2 | NVIDIA Cosmos vs V-JEPA 2 | Two green-index foundation-scale leaders with different views of world modeling. Cosmos emphasizes a platform for physical-AI simulation and generation, while V-JEPA 2 emphasizes self-supervised predictive representations for visual understanding and robot control. |
| PlayWorld vs V-JEPA 2 | PlayWorld vs V-JEPA 2 | Two green-index models pushing robotics-relevant world understanding in different ways. PlayWorld is a robot-play simulator for manipulation and policy improvement, while V-JEPA 2 is a self-supervised video predictor optimized for physical reasoning and zero-shot robot planning. |
Crawler-readable guide links tied to this model.
| Guide | Summary |
|---|---|
| World Models vs Large Language Models: A Practitioner's Guide | How world models differ from LLMs in objective, architecture and capability, and why both paradigms are likely to converge on the path to general-purpose AI. |
Connected research areas surfaced directly in static HTML.
| Topic | Summary |
|---|---|
| Video World Models | How video world models learn physics, temporal consistency, and interactive simulation from large-scale video, from Sora and Genie to Cosmos and V-JEPA. |
| World Model Evaluation | How to evaluate world models across rollout quality, benchmark performance, planning utility, and downstream transfer instead of relying on visual plausibility alone. |
FAQ answers rendered directly into static HTML for extractable responses.
V-JEPA 2 dramatically scales up the architecture and training data, achieving state-of-the-art visual understanding (surpassing supervised models) and demonstrating zero-shot robot control, capabilities V-JEPA 1 didn't exhibit.
Following LeCun's JEPA philosophy, predicting in latent space avoids wasting capacity on irrelevant pixel-level details (textures, exact colors) and focuses on learning meaningful physical representations.
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.
This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-04-10.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary model and lab sources embedded in static HTML.