New: the Timeline is live. Track world model releases, papers, and benchmark updates in real time.
world-models.io
The Knowledge Hub for AI World Models

Stable Video Diffusion

Stable Video Diffusion (SVD) is Stability AI's open-source foundation model for generative video, producing high-quality video from single image conditioning.

robotics model-based-rl simulation embodied-ai

Key Attributes

Static HTML snapshot of the model record for crawlers and no-JS readers.

AttributeValue
ModelStable Video Diffusion
Lab / OrganizationStability AI
CategoryGenerative World Model
SubtypeVideo Diffusion Model
World Model TypeImage-to-video diffusion model
Primary DomainVideo Generation
Architecture3D UNet latent diffusion with temporal attention layers
ModalityVisual (Image → Video)
Training MethodMulti-stage training: image pretraining → video fine-tuning on curated dataset
Statusactive
Year2023
Performance Index57/100 (high confidence, v1.1)

About Stable Video Diffusion

Main editorial body preserved directly in static HTML.

Stable Video Diffusion (SVD) is an open-source latent video diffusion model built upon the Stable Diffusion image model. It generates short, temporally coherent video sequences from a single conditioning image. SVD demonstrates strong motion dynamics and physical plausibility, making it a foundational building block for researchers working on video generation, world simulation, and controllable content creation. Its open weights have enabled rapid community adoption and downstream applications.

Stable Video Diffusion is an image-to-video diffusion model developed by Stability AI in 2023 for video generation.

Editorial Snapshot

Short extractable facts for answer engines and no-JS readers.

SignalValue
DefinitionStable Video Diffusion is an image-to-video diffusion model developed by Stability AI in 2023 for video generation.
Short DescriptionAn open-source foundation model for video generation, producing temporally consistent video from a single image.
Benchmark Rows1
FAQ Entries2
Related Models4
Related Guides0
Related Research Topics1
Last Updated2026-04-07

Notable Features

Key capabilities associated with this model.

  • Open-source weights
  • Image-to-video generation
  • Temporally consistent motion
  • Multi-stage curation pipeline for training data

Use Cases

Representative applications attached to this model record.

Creative video generationAnimation from stillsResearch on video dynamicsDownstream video editing

Strengths and Limitations

Balanced assessment surfaced in static HTML.

Strengths

  • Open-source and widely adopted
  • Strong temporal coherence
  • Good physical plausibility for short clips
  • Efficient latent space architecture

Limitations

  • Short video duration (2-4 seconds)
  • Limited resolution
  • No text-to-video (image conditioning only)
  • Occasional temporal artifacts

Benchmarks

Published benchmark evidence attached to this model record.

BenchmarkMetricResultSource
FVD (UCF-101)Fréchet Video DistanceCompetitive with closed-source modelsSource

References and Citations

Primary references preserved in static HTML for citation extraction.

ReferenceLink
Blattmann et al., 2023. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets. arXiv:2311.15127Open source

Related Models

Nearby models linked from the current editorial record.

ModelCategoryWorld Model TypeIndex v1.1
SoraGenerative World ModelText-to-video world simulator63/100
Genie 2Generative World ModelGenerative environment model79/100
Emu VideoGenerative World ModelFactorized text-to-video model49/100
NVIDIA CosmosFoundation World ModelVideo world foundation model87/100

Direct Comparisons

Side-by-side comparisons already connected to this model.

ComparisonMatchupSummary
Sora vs Gen-3 AlphaSora vs Gen-3 AlphaThe two leading commercial video generation models. Sora emphasizes physical world simulation and long-form coherence, while Gen-3 Alpha focuses on fine-grained creative control and production-ready tools.
Stable Video Diffusion vs Emu VideoStable Video Diffusion vs Emu VideoTwo image-to-video models: SVD is open-source and community-driven, while Emu Video is Meta's factorized approach that separates image and motion generation for better controllability.

Research Topics Referencing This Model

Connected research areas surfaced directly in static HTML.

TopicSummary
Diffusion World ModelsHow diffusion world models generate future states, preserve richer visual detail, and power video simulation systems such as DIAMOND, Sora, and Cosmos.

Frequently Asked Questions

FAQ answers rendered directly into static HTML for extractable responses.

How does SVD compare to Sora?

SVD is open-source and generates short (2-4s) videos from images. Sora is closed-source, generates longer videos from text, and shows stronger physical reasoning. SVD is more accessible for research.

Can SVD be used for world simulation?

While not designed as a world simulator, SVD's ability to generate physically plausible motion makes it a useful building block for world simulation research.

Quick Answer

Short extractable summary preserved directly in static HTML.

  • Stable Video Diffusion is an image-to-video diffusion model developed by Stability AI in 2023 for video generation.
  • Use this page when you need a fast read on how Stable Video Diffusion fits into the generative world model landscape, then validate the details in the benchmarks, citations, and related pages.
  • A key strength surfaced in the editorial record is open-source and widely adopted.

Editorial Trust Signals

Editorial provenance and refresh policy preserved directly in static HTML.

Published by world-models.io editorial board.

Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.

This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.

Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-04-07.

Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.

Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.

Primary sources onlyLast reviewed date visibleMethodology documentedSource links included

External Sources

Primary model and lab sources embedded in static HTML.

References

  1. [1] Blattmann et al., 2023. Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets. arXiv:2311.15127