New: the Timeline is live. Track world model releases, papers, and benchmark updates in real time.
world-models.io
The Knowledge Hub for AI World Models

V-JEPA

V-JEPA (Video Joint Embedding Predictive Architecture) learns visual world dynamics through self-supervised prediction in abstract representation space, without pixel reconstruction.

robotics model-based-rl simulation embodied-ai

Key Attributes

Static HTML snapshot of the model record for crawlers and no-JS readers.

AttributeValue
ModelV-JEPA
Lab / OrganizationMeta FAIR
CategorySelf-Supervised World Model
SubtypeJoint Embedding Predictive Architecture
World Model TypeSelf-supervised visual world model
Primary DomainVideo understanding
ArchitectureVision Transformer with joint embedding predictive objective
ModalityVideo
Training MethodSelf-supervised prediction in abstract representation space (no pixel reconstruction)
Statusactive
Year2024
Performance Index70/100 (medium confidence, v1.1)

About V-JEPA

Main editorial body preserved directly in static HTML.

V-JEPA follows Yann LeCun's JEPA framework to learn visual world models from video without pixel-level reconstruction. Instead of predicting pixels, it predicts abstract representations of future video frames, learning a world model that captures the causal structure of visual scenes. This approach avoids the pitfalls of pixel-level prediction while learning meaningful physical dynamics.

V-JEPA is a self-supervised visual world model developed by Meta in 2024 for video understanding.

Editorial Snapshot

Short extractable facts for answer engines and no-JS readers.

SignalValue
DefinitionV-JEPA is a self-supervised visual world model developed by Meta in 2024 for video understanding.
Short DescriptionVideo Joint Embedding Predictive Architecture: learns visual world models through self-supervised video prediction in abstract representation space.
Benchmark Rows1
FAQ Entries1
Related Models1
Related Guides1
Related Research Topics4
Last Updated2026-03-13

Notable Features

Key capabilities associated with this model.

  • No pixel-level reconstruction needed
  • Learns in abstract representation space
  • Follows LeCun's JEPA framework
  • Captures physical dynamics from video

Use Cases

Representative applications attached to this model record.

Video understandingPhysical reasoningVisual representation learningDownstream vision tasks

Strengths and Limitations

Balanced assessment surfaced in static HTML.

Strengths

  • No reconstruction loss artifacts
  • Learns meaningful physical dynamics
  • Strong transfer to downstream tasks
  • Theoretically principled (JEPA)

Limitations

  • Not yet used for RL policy learning
  • Abstract representations may miss fine details
  • Emerging approach, less battle-tested

Benchmarks

Published benchmark evidence attached to this model record.

BenchmarkMetricResultSource
Video Understanding TasksK400 Accuracy 81.3 %Competitive with supervised methodsSource

References and Citations

Primary references preserved in static HTML for citation extraction.

ReferenceLink
Bardes et al., 2024. V-JEPA: Video Joint Embedding Predictive Architecture.Open source

Related Models

Nearby models linked from the current editorial record.

ModelCategoryWorld Model TypeIndex v1.1
DreamerV3Model-Based RLImagination-based dynamics model88/100

Direct Comparisons

Side-by-side comparisons already connected to this model.

ComparisonMatchupSummary
World Models vs LLMsWorld Models vs Large Language ModelsWorld models and LLMs represent fundamentally different approaches to AI. World models learn causal dynamics of physical environments; LLMs learn statistical patterns over text. Both are essential for the future of AI.
V-JEPA vs Video Generation ModelsV-JEPA (Meta) vs Video Generation Models (Sora, Cosmos)V-JEPA and video generation models like Sora both learn from video, but follow opposite philosophies: V-JEPA predicts in abstract representation space without generating pixels, while video generation models focus on producing realistic pixel outputs.
I-JEPA vs MAE (Masked Autoencoders)I-JEPA (Meta FAIR) vs MAE (Meta / He et al.)I-JEPA and MAE are both self-supervised image learning methods, but they follow opposite philosophies: I-JEPA predicts in abstract representation space, while MAE reconstructs masked pixels.
V-JEPA vs I-JEPAV-JEPA vs I-JEPABoth implement Yann LeCun's JEPA framework for self-supervised learning, but V-JEPA operates on video (temporal dynamics) while I-JEPA operates on static images (spatial structure).
3D-VLA vs I-JEPA3D-VLA vs I-JEPATwo approaches to learning representations for embodied intelligence: 3D-VLA combines 3D perception with language-conditioned action planning, while I-JEPA learns abstract visual representations through self-supervised prediction in latent space.
V-JEPA vs NVIDIA CosmosV-JEPA vs NVIDIA CosmosTwo foundation-scale approaches to world understanding: V-JEPA learns predictive video representations through self-supervised masking, while Cosmos builds a full-stack world simulation platform for physical AI.
AMI vs Ha World ModelAMI vs Ha World ModelTwo pioneering cognitive-inspired world models: Ha's 2018 World Model introduced the VAE+RNN+Controller architecture, while AMI proposes an autonomous machine intelligence framework inspired by biological cognition.
LWM vs V-JEPALarge World Model (LWM) vs V-JEPATwo approaches to learning world understanding from video. LWM uses autoregressive prediction over million-length sequences, while V-JEPA predicts abstract latent representations without pixel reconstruction.

Guides Referencing This Model

Crawler-readable guide links tied to this model.

GuideSummary
How to Read World Models PapersA practical reading path through world-model research, from foundational concepts to latent dynamics, planning, simulators, and self-supervised approaches.

Research Topics Referencing This Model

Connected research areas surfaced directly in static HTML.

TopicSummary
Self-Supervised World ModelsHow self-supervised world models learn environment dynamics without rewards, from JEPA and V-JEPA to predictive latent representations.
World Models vs LLMsThe key differences between world models and LLMs across objective, architecture, planning, physical reasoning, and embodied AI use cases.
World Models: A Comprehensive SurveyA survey of AI world models covering taxonomy, leading architectures, landmark systems, open challenges, and future research directions.
Video World ModelsHow video world models learn physics, temporal consistency, and interactive simulation from large-scale video, from Sora and Genie to Cosmos and V-JEPA.

Timeline Mentions

Recent timeline events connected to this model.

EventPublishedSourceSummary
Meta FAIR releases V-JEPA 2 checkpoints under non-commercial license2026-04-08Meta FAIRMeta FAIR publishes V-JEPA 2 model checkpoints (ViT-L, ViT-H, ViT-g) on Hugging Face under a research-only license.
Meta FAIR publishes V-JEPA 2 paper: video prediction at scale without pixel reconstruction2026-03-10arXivMeta FAIR introduces V-JEPA 2, extending the Joint Embedding Predictive Architecture to video.

Frequently Asked Questions

FAQ answers rendered directly into static HTML for extractable responses.

How does V-JEPA differ from video generation models?

V-JEPA predicts in abstract representation space rather than generating pixels, learning meaningful causal dynamics without reconstruction artifacts.

Quick Answer

Short extractable summary preserved directly in static HTML.

  • V-JEPA is a self-supervised visual world model developed by Meta in 2024 for video understanding.
  • Use this page when you need a fast read on how V-JEPA fits into the self-supervised world model landscape, then validate the details in the benchmarks, citations, and related pages.
  • A key strength surfaced in the editorial record is no reconstruction loss artifacts.

Editorial Trust Signals

Editorial provenance and refresh policy preserved directly in static HTML.

Published by world-models.io editorial board.

Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.

This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.

Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-03-13.

Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.

Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.

Primary sources onlyLast reviewed date visibleMethodology documentedSource links included

External Sources

Primary model and lab sources embedded in static HTML.

References

  1. [1] Bardes et al., 2024. V-JEPA: Video Joint Embedding Predictive Architecture.