Static HTML snapshot of the model record for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Model | Genie |
| Lab / Organization | DeepMind |
| Category | Generative World Model |
| Subtype | Generative Interactive Environment |
| World Model Type | Action-controllable generative model |
| Primary Domain | 2D environment generation |
| Architecture | Video tokenizer (ST-ViViT) + Latent Action Model + Dynamics Model (MaskGIT-style) |
| Modality | Image → Interactive 2D Environment |
| Training Method | Unsupervised learning from 200K hours of internet video with no action labels |
| Status | foundational |
| Year | 2024 |
| Performance Index | 57/100 (medium confidence, v1.1) |
Main editorial body preserved directly in static HTML.
Genie (Generative Interactive Environment) is a foundation world model trained on 200K hours of unlabeled internet video. It learns latent actions from video alone and can generate playable 2D environments from a single image. The model consists of a video tokenizer, latent action model, and dynamics model, enabling interactive environment generation without any action labels during training.
Genie is an action-controllable generative model developed by Google DeepMind in 2024 for 2d environment generation.
Short extractable facts for answer engines and no-JS readers.
| Signal | Value |
|---|---|
| Definition | Genie is an action-controllable generative model developed by Google DeepMind in 2024 for 2d environment generation. |
| Short Description | The first generative interactive environment trained from unlabeled internet videos, capable of generating action-controllable 2D worlds. |
| Benchmark Rows | 1 |
| FAQ Entries | 1 |
| Related Models | 3 |
| Related Guides | 0 |
| Related Research Topics | 0 |
| Last Updated | 2026-03-14 |
Key capabilities associated with this model.
Representative applications attached to this model record.
Balanced assessment surfaced in static HTML.
Published benchmark evidence attached to this model record.
| Benchmark | Metric | Result | Source |
|---|---|---|---|
| Controllability | Action Accuracy 78.5 % | Strong action following | Source |
Primary references preserved in static HTML for citation extraction.
| Reference | Link |
|---|---|
| Bruce et al., 2024. Genie: Generative Interactive Environments. ICML 2024. | Open source |
Side-by-side comparisons already connected to this model.
| Comparison | Matchup | Summary |
|---|---|---|
| Genie vs Genie 2 | Genie (v1) vs Genie 2 | Genie pioneered unsupervised interactive environment generation from video. Genie 2 massively scales this approach to generate persistent, interactive 3D worlds from single images. |
| UniSim vs Genie 2 | UniSim vs Genie 2 | Both are large-scale generative world simulators, but UniSim focuses on unified simulation across real-world domains while Genie 2 generates persistent, explorable 3D environments from single images. |
| UniSim vs Genie 2 | UniSim vs Genie 2 | Two DeepMind generative world models targeting interactive simulation. UniSim uses a diffusion-based approach for universal simulation, while Genie 2 generates playable 3D environments from a single image. |
| STEVE-1 vs DreamerV3 | STEVE-1 vs DreamerV3 | Two approaches to open-world game AI. STEVE-1 uses video pre-training and instruction following, while DreamerV3 learns a world model from scratch via reinforcement learning. |
| Genie 3 vs Genie 2 | Genie 3 vs Genie 2 | Two generations of DeepMind's interactive world model. Genie 2 generates 3D environments from single images; Genie 3 generates them from text prompts in real time at 24fps with far greater diversity and consistency. |
FAQ answers rendered directly into static HTML for extractable responses.
Genie uses a Latent Action Model that discovers a discrete action space from video transitions alone, learning what actions could transform one frame into the next without any human annotation.
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.
This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-03-14.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary model and lab sources embedded in static HTML.