Static HTML snapshot of the model record for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Model | Large World Model (LWM) |
| Lab / Organization | UC Berkeley |
| Category | Foundation World Model |
| Subtype | Long-Context Multimodal Model |
| World Model Type | Million-length video-language world model |
| Primary Domain | Video Understanding |
| Architecture | LLaMA-based transformer with RingAttention for million-token context |
| Modality | Visual (Video) + Language |
| Training Method | Progressive context extension on interleaved video-text data |
| Status | active |
| Year | 2024 |
| Performance Index | 55/100 (medium confidence, v1.1) |
Main editorial body preserved directly in static HTML.
The Large World Model (LWM) extends the context length of multimodal transformers to over one million tokens, enabling processing of long videos interleaved with text. Built on the LLaMA architecture with RingAttention for efficient long-context training, LWM demonstrates that scaling sequence length (not just model size) unlocks emergent world understanding capabilities including long-horizon video comprehension, temporal reasoning, and cross-modal physical dynamics prediction.
Large World Model (LWM) is a million-length video-language world model developed by UC Berkeley in 2024 for video understanding.
Short extractable facts for answer engines and no-JS readers.
| Signal | Value |
|---|---|
| Definition | Large World Model (LWM) is a million-length video-language world model developed by UC Berkeley in 2024 for video understanding. |
| Short Description | A foundation model trained on 1M+ interleaved video and language tokens for long-horizon world understanding. |
| Benchmark Rows | 1 |
| FAQ Entries | 2 |
| Related Models | 4 |
| Related Guides | 0 |
| Related Research Topics | 0 |
| Last Updated | 2026-04-07 |
Key capabilities associated with this model.
Representative applications attached to this model record.
Balanced assessment surfaced in static HTML.
Published benchmark evidence attached to this model record.
| Benchmark | Metric | Result | Source |
|---|---|---|---|
| Long Video QA | Accuracy | State-of-the-art | Source |
Primary references preserved in static HTML for citation extraction.
| Reference | Link |
|---|---|
| Liu et al., 2024. World Model on Million-Length Video And Language With RingAttention. arXiv:2402.08855 | Open source |
Nearby models linked from the current editorial record.
| Model | Category | World Model Type | Index v1.1 |
|---|---|---|---|
| V-JEPA | Self-Supervised World Model | Self-supervised visual world model | 70/100 |
| Sora | Generative World Model | Text-to-video world simulator | 63/100 |
| NVIDIA Cosmos | Foundation World Model | Video world foundation model | 87/100 |
| AMI World Model | Foundation World Model | Multimodal generative world model | 38/100 |
Side-by-side comparisons already connected to this model.
| Comparison | Matchup | Summary |
|---|---|---|
| LWM vs V-JEPA | Large World Model (LWM) vs V-JEPA | Two approaches to learning world understanding from video. LWM uses autoregressive prediction over million-length sequences, while V-JEPA predicts abstract latent representations without pixel reconstruction. |
FAQ answers rendered directly into static HTML for extractable responses.
Longer context allows models to track persistent objects, understand cause-and-effect over time, and reason about physical dynamics across extended video sequences, essential capabilities for world understanding.
LWM uses RingAttention, a technique that distributes attention computation across multiple devices, enabling linear scaling of context length with available hardware.
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.
This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-04-07.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary model and lab sources embedded in static HTML.