New: the Timeline is live. Track world model releases, papers, and benchmark updates in real time.
world-models.io
The Knowledge Hub for AI World Models

V-JEPA 2

V-JEPA 2 (Video Joint Embedding Predictive Architecture 2) is Meta FAIR's self-supervised world model that achieves state-of-the-art visual understanding and enables zero-shot robot planning without pixel reconstruction. Fully open-source.

robotics model-based-rl simulation embodied-ai

Key Attributes

Static HTML snapshot of the model record for crawlers and no-JS readers.

AttributeValue
ModelV-JEPA 2
Lab / OrganizationMeta FAIR
CategorySelf-Supervised World Model
SubtypeVideo Prediction Model
World Model TypeJoint-embedding predictive world model for video understanding and robot planning
Primary DomainPhysical Reasoning & Robotics
ArchitectureVision Transformer with joint-embedding predictive architecture, latent space prediction
ModalityVideo → Latent Predictions
Training MethodSelf-supervised learning via latent space prediction from video, no pixel reconstruction
Statusactive
Year2025
Performance Index87/100 (medium confidence, v1.1)

About V-JEPA 2

Main editorial body preserved directly in static HTML.

V-JEPA 2 (Video Joint Embedding Predictive Architecture 2) is a self-supervised foundation world model from Meta FAIR that learns to understand, predict, and plan from video without relying on pixel-level reconstruction. By predicting in a learned latent space rather than pixel space, V-JEPA 2 avoids the computational overhead and irrelevant detail of generative approaches. It achieves state-of-the-art performance on visual understanding benchmarks and demonstrates zero-shot robot control capabilities, planning actions in new environments without task-specific fine-tuning. The model introduces new physical reasoning benchmarks and is fully open-source, marking a major step toward Yann LeCun's vision of world-model-based AI.

V-JEPA 2 is a joint-embedding predictive world model for video understanding and robot planning developed by Meta in 2025 for physical reasoning & robotics.

Editorial Snapshot

Short extractable facts for answer engines and no-JS readers.

SignalValue
DefinitionV-JEPA 2 is a joint-embedding predictive world model for video understanding and robot planning developed by Meta in 2025 for physical reasoning & robotics.
Short DescriptionMeta FAIR's self-supervised video world model achieving state-of-the-art visual understanding and enabling zero-shot robot control.
Benchmark Rows1
FAQ Entries2
Related Models3
Related Guides1
Related Research Topics2
Last Updated2026-04-10

Notable Features

Key capabilities associated with this model.

  • State-of-the-art visual understanding without pixel reconstruction
  • Zero-shot robot planning in unseen environments
  • New physical reasoning benchmarks
  • Fully open-source model and weights
  • Energy-efficient compared to generative world models

Use Cases

Representative applications attached to this model record.

Robot manipulation planningPhysical scene understandingVideo predictionEmbodied AIAction planning

Strengths and Limitations

Balanced assessment surfaced in static HTML.

Strengths

  • No pixel reconstruction overhead
  • Zero-shot transfer to robotics
  • Open-source
  • Strong physical reasoning
  • Scales efficiently

Limitations

  • Limited to visual modality
  • Requires large-scale video data
  • Robot experiments limited to manipulation tasks

Benchmarks

Published benchmark evidence attached to this model record.

BenchmarkMetricResultSource
Physical ReasoningPhyBench Accuracy 89.2 %State-of-the-artSource

References and Citations

Primary references preserved in static HTML for citation extraction.

ReferenceLink
Assran et al., 2025. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv:2506.09985.Open source

Related Models

Nearby models linked from the current editorial record.

ModelCategoryWorld Model TypeIndex v1.1
V-JEPASelf-Supervised World ModelSelf-supervised visual world model70/100
I-JEPASelf-Supervised World ModelSelf-supervised visual world model61/100
AMI World ModelFoundation World ModelMultimodal generative world model38/100

Direct Comparisons

Side-by-side comparisons already connected to this model.

ComparisonMatchupSummary
V-JEPA 2 vs V-JEPAV-JEPA 2 vs V-JEPAV-JEPA 2 dramatically scales up Meta FAIR's self-supervised video world model, achieving state-of-the-art visual understanding and zero-shot robot control, capabilities V-JEPA didn't demonstrate.
V-JEPA 2 vs I-JEPAV-JEPA 2 vs I-JEPATwo milestones of the JEPA roadmap: I-JEPA established self-supervised image representation by predicting in latent space; V-JEPA 2 extends the paradigm to video at foundation scale and demonstrates zero-shot robot control.
LeWorldModel vs DreamerV3LeWorldModel vs DreamerV3LeWorldModel revisits LeCun's energy-based JEPA philosophy for control, predicting in latent space without pixel reconstruction. DreamerV3 remains the canonical RSSM-based agent that learns by imagining pixel-grounded rollouts.
PlayWorld vs TD-MPC2PlayWorld vs TD-MPC2Two green-index models for robot decision-making, but with very different operating modes. PlayWorld learns a manipulation-focused world simulator from autonomous play, while TD-MPC2 combines latent dynamics with model-predictive control across a wide multi-task control benchmark suite.
Genie 3 vs NVIDIA CosmosGenie 3 vs NVIDIA CosmosTwo green-index frontier systems with different ambitions. Genie 3 is a real-time text-to-world interactive generator, while NVIDIA Cosmos is a broad physical-AI platform optimized for simulation infrastructure, robotics, and industrial world modeling.
Genie 3 vs V-JEPA 2Genie 3 vs V-JEPA 2Two green-index leaders that represent different frontier philosophies. Genie 3 is an interactive generative world model that turns text into playable environments, while V-JEPA 2 is a self-supervised latent predictor optimized for physical reasoning and zero-shot robot planning.
NVIDIA Cosmos vs V-JEPA 2NVIDIA Cosmos vs V-JEPA 2Two green-index foundation-scale leaders with different views of world modeling. Cosmos emphasizes a platform for physical-AI simulation and generation, while V-JEPA 2 emphasizes self-supervised predictive representations for visual understanding and robot control.
PlayWorld vs V-JEPA 2PlayWorld vs V-JEPA 2Two green-index models pushing robotics-relevant world understanding in different ways. PlayWorld is a robot-play simulator for manipulation and policy improvement, while V-JEPA 2 is a self-supervised video predictor optimized for physical reasoning and zero-shot robot planning.

Guides Referencing This Model

Crawler-readable guide links tied to this model.

GuideSummary
World Models vs Large Language Models: A Practitioner's GuideHow world models differ from LLMs in objective, architecture and capability, and why both paradigms are likely to converge on the path to general-purpose AI.

Research Topics Referencing This Model

Connected research areas surfaced directly in static HTML.

TopicSummary
Video World ModelsHow video world models learn physics, temporal consistency, and interactive simulation from large-scale video, from Sora and Genie to Cosmos and V-JEPA.
World Model EvaluationHow to evaluate world models across rollout quality, benchmark performance, planning utility, and downstream transfer instead of relying on visual plausibility alone.

Frequently Asked Questions

FAQ answers rendered directly into static HTML for extractable responses.

How does V-JEPA 2 differ from V-JEPA?

V-JEPA 2 dramatically scales up the architecture and training data, achieving state-of-the-art visual understanding (surpassing supervised models) and demonstrating zero-shot robot control, capabilities V-JEPA 1 didn't exhibit.

Why doesn't V-JEPA 2 use pixel reconstruction?

Following LeCun's JEPA philosophy, predicting in latent space avoids wasting capacity on irrelevant pixel-level details (textures, exact colors) and focuses on learning meaningful physical representations.

Quick Answer

Short extractable summary preserved directly in static HTML.

  • V-JEPA 2 is a joint-embedding predictive world model for video understanding and robot planning developed by Meta in 2025 for physical reasoning & robotics.
  • Use this page when you need a fast read on how V-JEPA 2 fits into the self-supervised world model landscape, then validate the details in the benchmarks, citations, and related pages.
  • A key strength surfaced in the editorial record is no pixel reconstruction overhead.

Editorial Trust Signals

Editorial provenance and refresh policy preserved directly in static HTML.

Published by world-models.io editorial board.

Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.

This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.

Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-04-10.

Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.

Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.

Primary sources onlyLast reviewed date visibleMethodology documentedSource links included

External Sources

Primary model and lab sources embedded in static HTML.

References

  1. [1] Assran et al., 2025. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. arXiv:2506.09985.