New: the Timeline is live. Track world model releases, papers, and benchmark updates in real time.
world-models.io
The Knowledge Hub for AI World Models

Video World Models

How video world models learn physics, temporal consistency, and interactive simulation from large-scale video, from Sora and Genie to Cosmos and V-JEPA.

robotics model-based-rl simulation embodied-ai

Research Snapshot

Static research summary generated from local editorial content.

AttributeValue
TopicVideo World Models
SummaryHow video world models learn physics, temporal consistency, and interactive simulation from large-scale video, from Sora and Genie to Cosmos and V-JEPA.
Related Models9
Citations3

What Are Video World Models?

Editorial body section preserved directly in static HTML.

Video world models learn environment dynamics directly from video sequences, treating video as a rich source of temporal, spatial, and physical structure. Instead of only modeling rewards or compact control states, they model how scenes evolve over time and how actions or prompts alter those futures.

Why Video Has Become a Core World Modeling Modality

Editorial body section preserved directly in static HTML.

Video captures object permanence, motion, occlusion, contact, and scene continuity at internet scale. That makes it one of the most information-dense training signals available for world models. As a result, video-based systems are increasingly central to robotics simulation, autonomous driving, embodied AI, and interactive environment generation.

Main Approaches: Generative, Predictive, and Interactive

Editorial body section preserved directly in static HTML.

The field now spans diffusion-based generators, autoregressive token models, JEPA-style predictive latent models, and action-conditioned interactive simulators. Sora and Cosmos emphasize generative realism, V-JEPA emphasizes predictive abstraction, and Genie-style systems push toward controllable environments rather than passive video synthesis.

Where Video World Models Are Strongest Today

Editorial body section preserved directly in static HTML.

Video world models are strongest when they need to produce realistic future frames, model plausible physical events, and create large quantities of synthetic training data. They are especially useful in driving simulation, environment generation, and embodied pretraining, where realism and temporal continuity matter.

Open Problems in Video World Modeling

Editorial body section preserved directly in static HTML.

The biggest open problems are long-horizon coherence, stable memory, action grounding, and evaluation. A model can generate visually plausible clips while still failing at causal control, object permanence, or physical consistency over longer interactions.

Related Models

ModelLabCategoryIndex v1.1
SoraOpenAIGenerative World Model63/100
Genie 2Google DeepMindGenerative World Model79/100
Genie 3Google DeepMindGenerative World Model89/100
NVIDIA CosmosNVIDIAFoundation World Model87/100
V-JEPAMetaSelf-Supervised World Model70/100
V-JEPA 2MetaSelf-Supervised World Model87/100
UniSimGoogle DeepMindGenerative World Model72/100
DIAMONDMicrosoft Research / University of GenevaModel-Based RL64/100
PandoraTsinghua University / ByteDanceGenerative World Model52/100

Frequently Asked Questions

FAQ answers rendered directly into static HTML for extractable responses.

Are video generation models the same as video world models?

Not always. A video generator can produce plausible sequences without being a strong interactive world model. Video world models become more useful when they preserve causality, memory, and controllability over future states.

Which systems define the video world model frontier?

Sora, Genie 3, NVIDIA Cosmos, V-JEPA 2, and UniSim each represent different parts of the frontier, from generative realism to predictive abstraction and interactive simulation.

Quick Answer

Short extractable summary preserved directly in static HTML.

  • Video World Models explains the core definition, methods, and systems involved in this research area.
  • This topic highlights the main trade-offs, open challenges, and practical implications for world models.
  • Related models and references connect the concept to concrete systems and primary sources.

Editorial Trust Signals

Editorial provenance and refresh policy preserved directly in static HTML.

Published by world-models.io editorial board.

Lead editor Bernard Grenat.

This research page curates topic explanations, linked models, and citations grounded in primary research sources.

Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-06-21.

Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.

Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.

Primary sources onlyLast reviewed date visibleMethodology documentedSource links included

External Sources

Primary research citations embedded in static HTML.

References

  1. [1] Brooks et al., 2024. Video Generation Models as World Simulators.
  2. [2] Bardes et al., 2024. V-JEPA: Video Joint Embedding Predictive Architecture.
  3. [3] Google DeepMind, 2025. Genie 3: A New Frontier for World Models.