Static HTML snapshot of the model record for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Model | Genie 2 |
| Lab / Organization | DeepMind |
| Category | Generative World Model |
| Subtype | Interactive Environment Generator |
| World Model Type | Generative environment model |
| Primary Domain | Environment generation |
| Architecture | Autoregressive latent diffusion transformer with action conditioning |
| Modality | Image → Interactive 3D Environment |
| Training Method | Large-scale video pre-training with action-conditioning |
| Status | active |
| Year | 2024 |
| Performance Index | 79/100 (medium confidence, v1.1) |
Main editorial body preserved directly in static HTML.
Genie 2 is a large-scale foundation world model capable of generating rich, interactive 3D environments. Given a single image prompt, it produces consistent, controllable worlds that maintain object permanence and realistic physics. The generated environments can be explored and interacted with, making them useful for training AI agents.
Genie 2 is a generative environment model developed by Google DeepMind in 2024 for environment generation.
Short extractable facts for answer engines and no-JS readers.
| Signal | Value |
|---|---|
| Definition | Genie 2 is a generative environment model developed by Google DeepMind in 2024 for environment generation. |
| Short Description | A foundation world model that generates diverse, playable 3D environments from a single image prompt. |
| Benchmark Rows | 1 |
| FAQ Entries | 1 |
| Related Models | 3 |
| Related Guides | 0 |
| Related Research Topics | 6 |
| Last Updated | 2026-03-14 |
Key capabilities associated with this model.
Representative applications attached to this model record.
Balanced assessment surfaced in static HTML.
Published benchmark evidence attached to this model record.
| Benchmark | Metric | Result | Source |
|---|---|---|---|
| Environment Consistency | Temporal Coherence 96.1 % | State-of-the-art | Source |
Primary references preserved in static HTML for citation extraction.
| Reference | Link |
|---|---|
| Google DeepMind, 2024. Genie 2: A Large-Scale Foundation World Model. | Open source |
Nearby models linked from the current editorial record.
| Model | Category | World Model Type | Index v1.1 |
|---|---|---|---|
| NVIDIA Cosmos | Foundation World Model | Video world foundation model | 87/100 |
| DreamerV3 | Model-Based RL | Imagination-based dynamics model | 88/100 |
| UniSim | Generative World Model | Action-conditioned video simulator | 72/100 |
Side-by-side comparisons already connected to this model.
| Comparison | Matchup | Summary |
|---|---|---|
| NVIDIA Cosmos vs DreamerV3 | NVIDIA Cosmos vs DreamerV3 | Cosmos and DreamerV3 represent two different scales and approaches to world modeling: Cosmos is a foundation-scale video world model platform for physical AI, while DreamerV3 is a sample-efficient RL agent with learned dynamics. |
| Sora vs Genie 2 | Sora (OpenAI) vs Genie 2 (DeepMind) | Sora and Genie 2 both generate video from prompts, but they approach world simulation very differently. Sora generates passive, high-fidelity videos from text; Genie 2 generates interactive, controllable 3D environments from images. |
| OASIS vs GameNGen | OASIS (Decart) vs GameNGen (Google Research) | OASIS and GameNGen both demonstrate neural networks functioning as real-time game engines, but they target different games and use different architectures. They represent the emerging frontier of neural game engines. |
| Genie 2 vs UniSim | Genie 2 vs UniSim | Both are generative world models that create interactive environments, but Genie 2 generates 3D worlds from single images while UniSim learns a universal action-conditioned simulator from diverse real-world data. |
| NVIDIA Cosmos vs Genie 2 | NVIDIA Cosmos vs Genie 2 | Two foundation-scale world models with different strategies: Cosmos is an open industrial platform for physical AI training, while Genie 2 is a DeepMind research system that generates interactive 3D environments from images. |
| OASIS vs DIAMOND | OASIS vs DIAMOND | Both use diffusion models as world models for interactive environments, but OASIS generates real-time playable Minecraft-like worlds while DIAMOND uses diffusion for model-based RL training in Atari. |
| Genie vs Genie 2 | Genie (v1) vs Genie 2 | Genie pioneered unsupervised interactive environment generation from video. Genie 2 massively scales this approach to generate persistent, interactive 3D worlds from single images. |
| Sora vs NVIDIA Cosmos | Sora (OpenAI) vs NVIDIA Cosmos | Both generate video from learned world dynamics, but Sora is a creative video generation model while Cosmos is an industrial platform for physical AI training and simulation. |
Connected research areas surfaced directly in static HTML.
| Topic | Summary |
|---|---|
| World Models vs LLMs | The key differences between world models and LLMs across objective, architecture, planning, physical reasoning, and embodied AI use cases. |
| Foundation World Models | How foundation world models such as Cosmos and Genie 2 bring large-scale learned simulation to robotics, autonomous driving, and physical AI. |
| AI Simulation Systems | How AI simulation systems and learned simulators reduce the reality gap and extend or replace hand-crafted engines for autonomous agents. |
| World Models: A Comprehensive Survey | A survey of AI world models covering taxonomy, leading architectures, landmark systems, open challenges, and future research directions. |
| Video World Models | How video world models learn physics, temporal consistency, and interactive simulation from large-scale video, from Sora and Genie to Cosmos and V-JEPA. |
| Diffusion World Models | How diffusion world models generate future states, preserve richer visual detail, and power video simulation systems such as DIAMOND, Sora, and Cosmos. |
Recent timeline events connected to this model.
| Event | Published | Source | Summary |
|---|---|---|---|
| Google DeepMind releases Genie 2: a foundation world model for 3D environments | 2026-03-14 | Google DeepMind | DeepMind announced Genie 2, a large-scale autoregressive world model capable of generating consistent... |
FAQ answers rendered directly into static HTML for extractable responses.
Genie 2 generates persistent 3D environments with consistent physics and object permanence, while Genie 1 was limited to 2D platformer-style worlds.
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.
This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-03-14.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary model and lab sources embedded in static HTML.