New: the Timeline is live. Track world model releases, papers, and benchmark updates in real time.
world-models.io
The Knowledge Hub for AI World Models

Frontier World Models 2026

A curated snapshot of the advanced world models shaping interactive simulation, self-supervised video understanding, and foundation-scale physical AI in 2026.

robotics model-based-rl simulation embodied-ai

Quick Answer

Short extractable summary preserved directly in static HTML.

  • A curated snapshot of the most advanced world models pushing the frontier in 2026, from interactive 3D generation and self-supervised video understanding to foundation-scale physical AI platforms.
  • The 2026 frontier is multipolar: interactive simulation, self-supervised physical understanding, and foundation-scale physical AI are advancing along different evaluation axes.

Why 2026 Matters

World models crossed a visible threshold in 2025-2026: Genie 3 demonstrated interactive 3D worlds at 24 fps and 720p with consistency over several minutes, V-JEPA 2 scaled self-supervised video understanding to 1.2B parameters with zero-shot robot transfer, and NVIDIA Cosmos shipped an open foundation platform for physical AI. These releases turned world models from research curiosity into deployable infrastructure.

Three Converging Tracks

The frontier organizes around three tracks. Interactive simulation (Genie 3, OASIS, GameNGen) generates controllable worlds in real time. Self-supervised understanding (V-JEPA 2) learns physics from raw video without labels. Foundation physical AI (NVIDIA Cosmos, AMI) trains general-purpose models on massive curated datasets for robotics and autonomous systems.

What Still Doesn't Work

Long-horizon coherence remains imperfect: scenes drift after the world memory window expires. Action conditioning is still coarse for video-native models like Sora. Sample-efficient model-based RL on real robots is improving but lags behind simulated benchmarks. Evaluation methodology has not yet caught up with foundation-scale models, leaving room for overclaims.

What to Watch Next

Convergence with LLMs is accelerating: hybrid systems pair world models for grounded prediction with language models for high-level reasoning. Open foundation models (Cosmos, V-JEPA 2) are accelerating the ecosystem. Expect specialized variants for surgery, manufacturing and household robotics within the next 12 months.

Frontier Model Directory

Curated model records rendered directly in static HTML.

ModelLab or companyCategoryWhy it matters
DreamerV3Google DeepMindModel-Based RLA general algorithm for mastering diverse domains with fixed hyperparameters through world model learning.
NVIDIA CosmosNVIDIAFoundation World ModelA platform of state-of-the-art generative world foundation models for physical AI development.
TD-MPC2MIT / MetaModel-Based RLA scalable world model agent that combines TD-learning with model-predictive control across 104 diverse tasks.
SoraOpenAIGenerative World ModelOpenAI's video generation model that simulates the physical world by generating realistic videos from text prompts.
DIAMONDMicrosoft Research / University of GenevaModel-Based RLDIffusion As a Model Of the eNvironment in Deep RL: uses diffusion models as world models for reinforcement learning agents.
OASISDecart / EtchedGenerative World ModelAn open-source real-time interactive world model that generates playable game environments at 20+ FPS entirely from a neural network.
GameNGenGoogle ResearchGenerative World ModelThe first neural model to simulate a complex game (DOOM) in real-time at high quality, making the game engine itself a neural network.
Genie 3Google DeepMindGenerative World ModelGoogle DeepMind's general-purpose world model that generates interactive 3D environments from text prompts in real time at 24fps.
V-JEPA 2MetaSelf-Supervised World ModelMeta FAIR's self-supervised video world model achieving state-of-the-art visual understanding and enabling zero-shot robot control.