Editorial definition preserved directly in static HTML.
The process of encoding raw video frames into a sequence of discrete or continuous tokens that a transformer can model. Video tokenizers (such as those released with NVIDIA Cosmos) are critical building blocks of foundation world models.
Video Tokenization is a glossary concept in the techniques layer of the world models knowledge base.
Static glossary definition snapshot for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Term | Video Tokenization |
| Category | Techniques |
| Definition | The process of encoding raw video frames into a sequence of discrete or continuous tokens that a transformer can model. Video tokenizers (such as those released with NVIDIA Cosmos) are critical building blocks of foundation world models. |
| Related Models | 2 |
| Related Research | 1 |
| Model | Lab | Category |
|---|---|---|
| NVIDIA Cosmos | NVIDIA | Foundation World Model |
| IRIS | Microsoft Research | Model-Based RL |
| Topic | Summary |
|---|---|
| Foundation World Models | How foundation world models such as Cosmos and Genie 2 bring large-scale learned simulation to robotics, autonomous driving, and physical AI. |
| Term | Category | Definition |
|---|---|---|
| Latent Space | Architecture | A compressed, abstract representation of data learned by a neural network. In world models, the latent space encodes environment states in a compact form that captures essential dynamics while discarding irrelevant details. Models like DreamerV3 and PlaNet operate entirely in latent space for efficient planning. |
| Foundation Model | Paradigms | A large-scale model trained on broad data that can be adapted to many downstream tasks. Foundation world models (like NVIDIA Cosmos and Genie 2) learn general-purpose representations of world dynamics, analogous to how GPT models serve as foundations for language tasks. |
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Bernard Grenat.
This glossary page publishes stable definitions linked to related models, research topics, and primary-source context.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-06-21.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary sources related to this term, surfaced directly in static HTML.