Static HTML snapshot of the model record for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Model | Emu Video |
| Lab / Organization | Meta FAIR |
| Category | Generative World Model |
| Subtype | Video Generation Model |
| World Model Type | Factorized text-to-video model |
| Primary Domain | Video generation |
| Architecture | Two-stage diffusion model: text-to-image + image-to-video |
| Modality | Text → Image → Video |
| Training Method | Factorized training: image diffusion + video diffusion with image conditioning |
| Status | active |
| Year | 2023 |
| Performance Index | 49/100 (medium confidence, v1.1) |
Main editorial body preserved directly in static HTML.
Emu Video simplifies text-to-video generation by factorizing it into two steps: text-to-image generation followed by image-to-video animation. This factored approach reduces the complexity of direct text-to-video generation while producing high-quality results. The model uses a diffusion-based architecture and achieves strong results compared to commercial video generation systems.
Emu Video is a factorized text-to-video model developed by Meta in 2023 for video generation.
Short extractable facts for answer engines and no-JS readers.
| Signal | Value |
|---|---|
| Definition | Emu Video is a factorized text-to-video model developed by Meta in 2023 for video generation. |
| Short Description | Meta's efficient video generation model using a factorized approach: first generate an image, then animate it into a video. |
| Benchmark Rows | 1 |
| FAQ Entries | 1 |
| Related Models | 2 |
| Related Guides | 0 |
| Related Research Topics | 0 |
| Last Updated | 2026-03-05 |
Key capabilities associated with this model.
Representative applications attached to this model record.
Balanced assessment surfaced in static HTML.
Published benchmark evidence attached to this model record.
| Benchmark | Metric | Result | Source |
|---|---|---|---|
| Human Evaluation | Human Preference 81 % preference | Preferred over competitors | Source |
Primary references preserved in static HTML for citation extraction.
| Reference | Link |
|---|---|
| Girdhar et al., 2023. Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning. arXiv:2311.10709 | Open source |
Nearby models linked from the current editorial record.
| Model | Category | World Model Type | Index v1.1 |
|---|---|---|---|
| Sora | Generative World Model | Text-to-video world simulator | 63/100 |
| NVIDIA Cosmos | Foundation World Model | Video world foundation model | 87/100 |
Side-by-side comparisons already connected to this model.
| Comparison | Matchup | Summary |
|---|---|---|
| Emu Video vs Sora | Emu Video vs Sora | Both are frontier video generation models, but with different ambitions: Emu Video focuses on efficient, high-quality short-form generation, while Sora pushes toward long-form, physically coherent world simulation. |
| Sora vs Emu Video | Sora vs Emu Video | Two generative video models from competing labs: Sora represents OpenAI's vision of video as world simulation, while Emu Video is Meta's efficient factorized approach to high-quality text-to-video generation. |
| Stable Video Diffusion vs Emu Video | Stable Video Diffusion vs Emu Video | Two image-to-video models: SVD is open-source and community-driven, while Emu Video is Meta's factorized approach that separates image and motion generation for better controllability. |
| PixVerse R1 vs Sora | PixVerse R1 vs Sora | PixVerse R1 introduces reasoning-trained generation to text-to-video, optimizing for prompt adherence and physical plausibility. Sora remains the reference for cinematic length and visual fidelity. |
FAQ answers rendered directly into static HTML for extractable responses.
Emu Video is primarily a video generation model. While it implicitly learns some physical dynamics, it lacks action conditioning and interactive capabilities that define true world models.
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.
This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-03-05.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary model and lab sources embedded in static HTML.