New: the Timeline is live. Track world model releases, papers, and benchmark updates in real time.
world-models.io
The Knowledge Hub for AI World Models

Genie

Genie (Generative Interactive Environment) is Google DeepMind's foundation world model trained on 200K hours of unlabeled internet video. It discovers latent actions automatically and generates playable 2D environments from a single image.

robotics model-based-rl simulation embodied-ai

Key Attributes

Static HTML snapshot of the model record for crawlers and no-JS readers.

AttributeValue
ModelGenie
Lab / OrganizationDeepMind
CategoryGenerative World Model
SubtypeGenerative Interactive Environment
World Model TypeAction-controllable generative model
Primary Domain2D environment generation
ArchitectureVideo tokenizer (ST-ViViT) + Latent Action Model + Dynamics Model (MaskGIT-style)
ModalityImage → Interactive 2D Environment
Training MethodUnsupervised learning from 200K hours of internet video with no action labels
Statusfoundational
Year2024
Performance Index57/100 (medium confidence, v1.1)

About Genie

Main editorial body preserved directly in static HTML.

Genie (Generative Interactive Environment) is a foundation world model trained on 200K hours of unlabeled internet video. It learns latent actions from video alone and can generate playable 2D environments from a single image. The model consists of a video tokenizer, latent action model, and dynamics model, enabling interactive environment generation without any action labels during training.

Genie is an action-controllable generative model developed by Google DeepMind in 2024 for 2d environment generation.

Editorial Snapshot

Short extractable facts for answer engines and no-JS readers.

SignalValue
DefinitionGenie is an action-controllable generative model developed by Google DeepMind in 2024 for 2d environment generation.
Short DescriptionThe first generative interactive environment trained from unlabeled internet videos, capable of generating action-controllable 2D worlds.
Benchmark Rows1
FAQ Entries1
Related Models3
Related Guides0
Related Research Topics0
Last Updated2026-03-14

Notable Features

Key capabilities associated with this model.

  • Trained on unlabeled internet video
  • Discovers latent action space automatically
  • Single image to playable world
  • 11B parameter foundation model

Use Cases

Representative applications attached to this model record.

2D environment generationAI agent trainingCreative game designWorld model research

Strengths and Limitations

Balanced assessment surfaced in static HTML.

Strengths

  • No action labels required
  • Massive scale training
  • Generative and interactive
  • Foundation for Genie 2

Limitations

  • 2D environments only
  • Low resolution output
  • Short generation horizons
  • Not publicly available

Benchmarks

Published benchmark evidence attached to this model record.

BenchmarkMetricResultSource
ControllabilityAction Accuracy 78.5 %Strong action followingSource

References and Citations

Primary references preserved in static HTML for citation extraction.

ReferenceLink
Bruce et al., 2024. Genie: Generative Interactive Environments. ICML 2024.Open source

Related Models

Nearby models linked from the current editorial record.

ModelCategoryWorld Model TypeIndex v1.1
Genie 2Generative World ModelGenerative environment model79/100
SoraGenerative World ModelText-to-video world simulator63/100
OASISGenerative World ModelReal-time playable world model66/100

Direct Comparisons

Side-by-side comparisons already connected to this model.

ComparisonMatchupSummary
Genie vs Genie 2Genie (v1) vs Genie 2Genie pioneered unsupervised interactive environment generation from video. Genie 2 massively scales this approach to generate persistent, interactive 3D worlds from single images.
UniSim vs Genie 2UniSim vs Genie 2Both are large-scale generative world simulators, but UniSim focuses on unified simulation across real-world domains while Genie 2 generates persistent, explorable 3D environments from single images.
UniSim vs Genie 2UniSim vs Genie 2Two DeepMind generative world models targeting interactive simulation. UniSim uses a diffusion-based approach for universal simulation, while Genie 2 generates playable 3D environments from a single image.
STEVE-1 vs DreamerV3STEVE-1 vs DreamerV3Two approaches to open-world game AI. STEVE-1 uses video pre-training and instruction following, while DreamerV3 learns a world model from scratch via reinforcement learning.
Genie 3 vs Genie 2Genie 3 vs Genie 2Two generations of DeepMind's interactive world model. Genie 2 generates 3D environments from single images; Genie 3 generates them from text prompts in real time at 24fps with far greater diversity and consistency.

Frequently Asked Questions

FAQ answers rendered directly into static HTML for extractable responses.

How does Genie learn actions without labels?

Genie uses a Latent Action Model that discovers a discrete action space from video transitions alone, learning what actions could transform one frame into the next without any human annotation.

Quick Answer

Short extractable summary preserved directly in static HTML.

  • Genie is an action-controllable generative model developed by Google DeepMind in 2024 for 2d environment generation.
  • Use this page when you need a fast read on how Genie fits into the generative world model landscape, then validate the details in the benchmarks, citations, and related pages.
  • A key strength surfaced in the editorial record is no action labels required.

Editorial Trust Signals

Editorial provenance and refresh policy preserved directly in static HTML.

Published by world-models.io editorial board.

Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.

This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.

Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-03-14.

Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.

Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.

Primary sources onlyLast reviewed date visibleMethodology documentedSource links included

External Sources

Primary model and lab sources embedded in static HTML.

References

  1. [1] Bruce et al., 2024. Genie: Generative Interactive Environments. ICML 2024.