New: the Timeline is live. Track world model releases, papers, and benchmark updates in real time.
world-models.io
The Knowledge Hub for AI World Models

Genie 2

Genie 2 is Google DeepMind's foundation world model that generates consistent, playable 3D environments from single images, demonstrating emergent physical understanding.

robotics model-based-rl simulation embodied-ai

Key Attributes

Static HTML snapshot of the model record for crawlers and no-JS readers.

AttributeValue
ModelGenie 2
Lab / OrganizationDeepMind
CategoryGenerative World Model
SubtypeInteractive Environment Generator
World Model TypeGenerative environment model
Primary DomainEnvironment generation
ArchitectureAutoregressive latent diffusion transformer with action conditioning
ModalityImage → Interactive 3D Environment
Training MethodLarge-scale video pre-training with action-conditioning
Statusactive
Year2024
Performance Index79/100 (medium confidence, v1.1)

About Genie 2

Main editorial body preserved directly in static HTML.

Genie 2 is a large-scale foundation world model capable of generating rich, interactive 3D environments. Given a single image prompt, it produces consistent, controllable worlds that maintain object permanence and realistic physics. The generated environments can be explored and interacted with, making them useful for training AI agents.

Genie 2 is a generative environment model developed by Google DeepMind in 2024 for environment generation.

Editorial Snapshot

Short extractable facts for answer engines and no-JS readers.

SignalValue
DefinitionGenie 2 is a generative environment model developed by Google DeepMind in 2024 for environment generation.
Short DescriptionA foundation world model that generates diverse, playable 3D environments from a single image prompt.
Benchmark Rows1
FAQ Entries1
Related Models3
Related Guides0
Related Research Topics6
Last Updated2026-03-14

Notable Features

Key capabilities associated with this model.

  • Single image to full 3D environment
  • Maintains object permanence
  • Action-controllable generation
  • Consistent physics simulation

Use Cases

Representative applications attached to this model record.

AI agent training environmentsGame prototypingSimulation generationEmbodied AI research

Strengths and Limitations

Balanced assessment surfaced in static HTML.

Strengths

  • Single image to full environment
  • Consistent physics
  • Action-controllable
  • Scalable generation

Limitations

  • Not publicly available
  • Limited to short generation horizons
  • Proprietary

Benchmarks

Published benchmark evidence attached to this model record.

BenchmarkMetricResultSource
Environment ConsistencyTemporal Coherence 96.1 %State-of-the-artSource

References and Citations

Primary references preserved in static HTML for citation extraction.

ReferenceLink
Google DeepMind, 2024. Genie 2: A Large-Scale Foundation World Model.Open source

Related Models

Nearby models linked from the current editorial record.

ModelCategoryWorld Model TypeIndex v1.1
NVIDIA CosmosFoundation World ModelVideo world foundation model87/100
DreamerV3Model-Based RLImagination-based dynamics model88/100
UniSimGenerative World ModelAction-conditioned video simulator72/100

Direct Comparisons

Side-by-side comparisons already connected to this model.

ComparisonMatchupSummary
NVIDIA Cosmos vs DreamerV3NVIDIA Cosmos vs DreamerV3Cosmos and DreamerV3 represent two different scales and approaches to world modeling: Cosmos is a foundation-scale video world model platform for physical AI, while DreamerV3 is a sample-efficient RL agent with learned dynamics.
Sora vs Genie 2Sora (OpenAI) vs Genie 2 (DeepMind)Sora and Genie 2 both generate video from prompts, but they approach world simulation very differently. Sora generates passive, high-fidelity videos from text; Genie 2 generates interactive, controllable 3D environments from images.
OASIS vs GameNGenOASIS (Decart) vs GameNGen (Google Research)OASIS and GameNGen both demonstrate neural networks functioning as real-time game engines, but they target different games and use different architectures. They represent the emerging frontier of neural game engines.
Genie 2 vs UniSimGenie 2 vs UniSimBoth are generative world models that create interactive environments, but Genie 2 generates 3D worlds from single images while UniSim learns a universal action-conditioned simulator from diverse real-world data.
NVIDIA Cosmos vs Genie 2NVIDIA Cosmos vs Genie 2Two foundation-scale world models with different strategies: Cosmos is an open industrial platform for physical AI training, while Genie 2 is a DeepMind research system that generates interactive 3D environments from images.
OASIS vs DIAMONDOASIS vs DIAMONDBoth use diffusion models as world models for interactive environments, but OASIS generates real-time playable Minecraft-like worlds while DIAMOND uses diffusion for model-based RL training in Atari.
Genie vs Genie 2Genie (v1) vs Genie 2Genie pioneered unsupervised interactive environment generation from video. Genie 2 massively scales this approach to generate persistent, interactive 3D worlds from single images.
Sora vs NVIDIA CosmosSora (OpenAI) vs NVIDIA CosmosBoth generate video from learned world dynamics, but Sora is a creative video generation model while Cosmos is an industrial platform for physical AI training and simulation.

Research Topics Referencing This Model

Connected research areas surfaced directly in static HTML.

TopicSummary
World Models vs LLMsThe key differences between world models and LLMs across objective, architecture, planning, physical reasoning, and embodied AI use cases.
Foundation World ModelsHow foundation world models such as Cosmos and Genie 2 bring large-scale learned simulation to robotics, autonomous driving, and physical AI.
AI Simulation SystemsHow AI simulation systems and learned simulators reduce the reality gap and extend or replace hand-crafted engines for autonomous agents.
World Models: A Comprehensive SurveyA survey of AI world models covering taxonomy, leading architectures, landmark systems, open challenges, and future research directions.
Video World ModelsHow video world models learn physics, temporal consistency, and interactive simulation from large-scale video, from Sora and Genie to Cosmos and V-JEPA.
Diffusion World ModelsHow diffusion world models generate future states, preserve richer visual detail, and power video simulation systems such as DIAMOND, Sora, and Cosmos.

Timeline Mentions

Recent timeline events connected to this model.

EventPublishedSourceSummary
Google DeepMind releases Genie 2: a foundation world model for 3D environments2026-03-14Google DeepMindDeepMind announced Genie 2, a large-scale autoregressive world model capable of generating consistent...

Frequently Asked Questions

FAQ answers rendered directly into static HTML for extractable responses.

How does Genie 2 differ from Genie 1?

Genie 2 generates persistent 3D environments with consistent physics and object permanence, while Genie 1 was limited to 2D platformer-style worlds.

Quick Answer

Short extractable summary preserved directly in static HTML.

  • Genie 2 is a Generative environment model developed by Google DeepMind in 2024 for environment generation.
  • Use this page when you need a fast read on how Genie 2 fits into the generative world model landscape, then validate the details in the benchmarks, citations, and related pages.
  • A key strength surfaced in the editorial record is single image to full environment.

Editorial Trust Signals

Editorial provenance and refresh policy preserved directly in static HTML.

Published by world-models.io editorial board.

Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.

This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.

Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-03-14.

Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.

Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.

Primary sources onlyLast reviewed date visibleMethodology documentedSource links included

External Sources

Primary model and lab sources embedded in static HTML.

References

  1. [1] Google DeepMind, 2024. Genie 2: A Large-Scale Foundation World Model.