New: the Timeline is live. Track world model releases, papers, and benchmark updates in real time.
world-models.io
The Knowledge Hub for AI World Models

Emu Video

Emu Video is Meta's video generation model that uses a factorized approach: generating a key image from text, then animating it into a video sequence.

robotics model-based-rl simulation embodied-ai

Key Attributes

Static HTML snapshot of the model record for crawlers and no-JS readers.

AttributeValue
ModelEmu Video
Lab / OrganizationMeta FAIR
CategoryGenerative World Model
SubtypeVideo Generation Model
World Model TypeFactorized text-to-video model
Primary DomainVideo generation
ArchitectureTwo-stage diffusion model: text-to-image + image-to-video
ModalityText → Image → Video
Training MethodFactorized training: image diffusion + video diffusion with image conditioning
Statusactive
Year2023
Performance Index49/100 (medium confidence, v1.1)

About Emu Video

Main editorial body preserved directly in static HTML.

Emu Video simplifies text-to-video generation by factorizing it into two steps: text-to-image generation followed by image-to-video animation. This factored approach reduces the complexity of direct text-to-video generation while producing high-quality results. The model uses a diffusion-based architecture and achieves strong results compared to commercial video generation systems.

Emu Video is a factorized text-to-video model developed by Meta in 2023 for video generation.

Editorial Snapshot

Short extractable facts for answer engines and no-JS readers.

SignalValue
DefinitionEmu Video is a factorized text-to-video model developed by Meta in 2023 for video generation.
Short DescriptionMeta's efficient video generation model using a factorized approach: first generate an image, then animate it into a video.
Benchmark Rows1
FAQ Entries1
Related Models2
Related Guides0
Related Research Topics0
Last Updated2026-03-05

Notable Features

Key capabilities associated with this model.

  • Factorized two-stage approach
  • High visual quality
  • Simpler than end-to-end methods
  • Strong human evaluation results

Use Cases

Representative applications attached to this model record.

Video content creationAnimation from imagesCreative toolsVisual storytelling

Strengths and Limitations

Balanced assessment surfaced in static HTML.

Strengths

  • Simplified architecture
  • High quality output
  • Efficient factored approach
  • Strong human preference scores

Limitations

  • Not action-conditioned
  • Short video clips only
  • No interactive control

Benchmarks

Published benchmark evidence attached to this model record.

BenchmarkMetricResultSource
Human EvaluationHuman Preference 81 % preferencePreferred over competitorsSource

References and Citations

Primary references preserved in static HTML for citation extraction.

ReferenceLink
Girdhar et al., 2023. Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning. arXiv:2311.10709Open source

Related Models

Nearby models linked from the current editorial record.

ModelCategoryWorld Model TypeIndex v1.1
SoraGenerative World ModelText-to-video world simulator63/100
NVIDIA CosmosFoundation World ModelVideo world foundation model87/100

Direct Comparisons

Side-by-side comparisons already connected to this model.

ComparisonMatchupSummary
Emu Video vs SoraEmu Video vs SoraBoth are frontier video generation models, but with different ambitions: Emu Video focuses on efficient, high-quality short-form generation, while Sora pushes toward long-form, physically coherent world simulation.
Sora vs Emu VideoSora vs Emu VideoTwo generative video models from competing labs: Sora represents OpenAI's vision of video as world simulation, while Emu Video is Meta's efficient factorized approach to high-quality text-to-video generation.
Stable Video Diffusion vs Emu VideoStable Video Diffusion vs Emu VideoTwo image-to-video models: SVD is open-source and community-driven, while Emu Video is Meta's factorized approach that separates image and motion generation for better controllability.
PixVerse R1 vs SoraPixVerse R1 vs SoraPixVerse R1 introduces reasoning-trained generation to text-to-video, optimizing for prompt adherence and physical plausibility. Sora remains the reference for cinematic length and visual fidelity.

Frequently Asked Questions

FAQ answers rendered directly into static HTML for extractable responses.

Is Emu Video a world model?

Emu Video is primarily a video generation model. While it implicitly learns some physical dynamics, it lacks action conditioning and interactive capabilities that define true world models.

Quick Answer

Short extractable summary preserved directly in static HTML.

  • Emu Video is a factorized text-to-video model developed by Meta in 2023 for video generation.
  • Use this page when you need a fast read on how Emu Video fits into the generative world model landscape, then validate the details in the benchmarks, citations, and related pages.
  • A key strength surfaced in the editorial record is simplified architecture.

Editorial Trust Signals

Editorial provenance and refresh policy preserved directly in static HTML.

Published by world-models.io editorial board.

Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.

This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.

Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-03-05.

Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.

Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.

Primary sources onlyLast reviewed date visibleMethodology documentedSource links included

External Sources

Primary model and lab sources embedded in static HTML.

References

  1. [1] Girdhar et al., 2023. Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning. arXiv:2311.10709