Short extractable summary preserved directly in static HTML.
World models crossed a visible threshold in 2025-2026: Genie 3 demonstrated interactive 3D worlds at 24 fps and 720p with consistency over several minutes, V-JEPA 2 scaled self-supervised video understanding to 1.2B parameters with zero-shot robot transfer, and NVIDIA Cosmos shipped an open foundation platform for physical AI. These releases turned world models from research curiosity into deployable infrastructure.
The frontier organizes around three tracks. Interactive simulation (Genie 3, OASIS, GameNGen) generates controllable worlds in real time. Self-supervised understanding (V-JEPA 2) learns physics from raw video without labels. Foundation physical AI (NVIDIA Cosmos, AMI) trains general-purpose models on massive curated datasets for robotics and autonomous systems.
Long-horizon coherence remains imperfect: scenes drift after the world memory window expires. Action conditioning is still coarse for video-native models like Sora. Sample-efficient model-based RL on real robots is improving but lags behind simulated benchmarks. Evaluation methodology has not yet caught up with foundation-scale models, leaving room for overclaims.
Convergence with LLMs is accelerating: hybrid systems pair world models for grounded prediction with language models for high-level reasoning. Open foundation models (Cosmos, V-JEPA 2) are accelerating the ecosystem. Expect specialized variants for surgery, manufacturing and household robotics within the next 12 months.
Curated model records rendered directly in static HTML.
| Model | Lab or company | Category | Why it matters |
|---|---|---|---|
| DreamerV3 | Google DeepMind | Model-Based RL | A general algorithm for mastering diverse domains with fixed hyperparameters through world model learning. |
| NVIDIA Cosmos | NVIDIA | Foundation World Model | A platform of state-of-the-art generative world foundation models for physical AI development. |
| TD-MPC2 | MIT / Meta | Model-Based RL | A scalable world model agent that combines TD-learning with model-predictive control across 104 diverse tasks. |
| Sora | OpenAI | Generative World Model | OpenAI's video generation model that simulates the physical world by generating realistic videos from text prompts. |
| DIAMOND | Microsoft Research / University of Geneva | Model-Based RL | DIffusion As a Model Of the eNvironment in Deep RL: uses diffusion models as world models for reinforcement learning agents. |
| OASIS | Decart / Etched | Generative World Model | An open-source real-time interactive world model that generates playable game environments at 20+ FPS entirely from a neural network. |
| GameNGen | Google Research | Generative World Model | The first neural model to simulate a complex game (DOOM) in real-time at high quality, making the game engine itself a neural network. |
| Genie 3 | Google DeepMind | Generative World Model | Google DeepMind's general-purpose world model that generates interactive 3D environments from text prompts in real time at 24fps. |
| V-JEPA 2 | Meta | Self-Supervised World Model | Meta FAIR's self-supervised video world model achieving state-of-the-art visual understanding and enabling zero-shot robot control. |