Static HTML snapshot of the model record for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Model | Stable Video Diffusion |
| Lab / Organization | Stability AI |
| Category | Generative World Model |
| Subtype | Video Diffusion Model |
| World Model Type | Image-to-video diffusion model |
| Primary Domain | Video Generation |
| Architecture | 3D UNet latent diffusion with temporal attention layers |
| Modality | Visual (Image → Video) |
| Training Method | Multi-stage training: image pretraining → video fine-tuning on curated dataset |
| Status | active |
| Year | 2023 |
| Performance Index | 57/100 (high confidence, v1.1) |
Main editorial body preserved directly in static HTML.
Stable Video Diffusion (SVD) is an open-source latent video diffusion model built upon the Stable Diffusion image model. It generates short, temporally coherent video sequences from a single conditioning image. SVD demonstrates strong motion dynamics and physical plausibility, making it a foundational building block for researchers working on video generation, world simulation, and controllable content creation. Its open weights have enabled rapid community adoption and downstream applications.
Stable Video Diffusion is an image-to-video diffusion model developed by Stability AI in 2023 for video generation.
Short extractable facts for answer engines and no-JS readers.
| Signal | Value |
|---|---|
| Definition | Stable Video Diffusion is an image-to-video diffusion model developed by Stability AI in 2023 for video generation. |
| Short Description | An open-source foundation model for video generation, producing temporally consistent video from a single image. |
| Benchmark Rows | 1 |
| FAQ Entries | 2 |
| Related Models | 4 |
| Related Guides | 0 |
| Related Research Topics | 1 |
| Last Updated | 2026-04-07 |
Key capabilities associated with this model.
Representative applications attached to this model record.
Balanced assessment surfaced in static HTML.
Published benchmark evidence attached to this model record.
| Benchmark | Metric | Result | Source |
|---|---|---|---|
| FVD (UCF-101) | Fréchet Video Distance | Competitive with closed-source models | Source |
Primary references preserved in static HTML for citation extraction.
| Reference | Link |
|---|---|
| Blattmann et al., 2023. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets. arXiv:2311.15127 | Open source |
Nearby models linked from the current editorial record.
| Model | Category | World Model Type | Index v1.1 |
|---|---|---|---|
| Sora | Generative World Model | Text-to-video world simulator | 63/100 |
| Genie 2 | Generative World Model | Generative environment model | 79/100 |
| Emu Video | Generative World Model | Factorized text-to-video model | 49/100 |
| NVIDIA Cosmos | Foundation World Model | Video world foundation model | 87/100 |
Side-by-side comparisons already connected to this model.
| Comparison | Matchup | Summary |
|---|---|---|
| Sora vs Gen-3 Alpha | Sora vs Gen-3 Alpha | The two leading commercial video generation models. Sora emphasizes physical world simulation and long-form coherence, while Gen-3 Alpha focuses on fine-grained creative control and production-ready tools. |
| Stable Video Diffusion vs Emu Video | Stable Video Diffusion vs Emu Video | Two image-to-video models: SVD is open-source and community-driven, while Emu Video is Meta's factorized approach that separates image and motion generation for better controllability. |
Connected research areas surfaced directly in static HTML.
| Topic | Summary |
|---|---|
| Diffusion World Models | How diffusion world models generate future states, preserve richer visual detail, and power video simulation systems such as DIAMOND, Sora, and Cosmos. |
FAQ answers rendered directly into static HTML for extractable responses.
SVD is open-source and generates short (2-4s) videos from images. Sora is closed-source, generates longer videos from text, and shows stronger physical reasoning. SVD is more accessible for research.
While not designed as a world simulator, SVD's ability to generate physically plausible motion makes it a useful building block for world simulation research.
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.
This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-04-07.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary model and lab sources embedded in static HTML.