Static HTML snapshot of the model record for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Model | 3D-VLA |
| Lab / Organization | MIT |
| Category | Foundation World Model |
| Subtype | 3D Vision-Language-Action Model |
| World Model Type | 3D-aware embodied world model |
| Primary Domain | Robotics / Embodied AI |
| Architecture | 3D-aware vision-language model with integrated world model and action decoder |
| Modality | 3D Point Clouds + Language + Actions |
| Training Method | Multi-task training: 3D understanding, world modeling, and action generation |
| Status | emerging |
| Year | 2024 |
| Performance Index | 47/100 (low confidence, v1.1) |
Main editorial body preserved directly in static HTML.
3D-VLA integrates 3D perception, language understanding, and action generation with an explicit world model component. The model can reason about 3D scenes, predict future states, and generate robot actions, all within a unified architecture. By incorporating a world model, 3D-VLA can plan ahead by imagining consequences of actions in 3D space before executing them.
3D-VLA is a 3d-aware embodied world model developed by MIT / Tsinghua in 2024 for robotics / embodied ai.
Short extractable facts for answer engines and no-JS readers.
| Signal | Value |
|---|---|
| Definition | 3D-VLA is a 3d-aware embodied world model developed by MIT / Tsinghua in 2024 for robotics / embodied ai. |
| Short Description | A 3D vision-language-action model with a built-in world model for embodied AI, enabling 3D-aware reasoning, planning, and action generation. |
| Benchmark Rows | 0 |
| FAQ Entries | 1 |
| Related Models | 3 |
| Related Guides | 0 |
| Related Research Topics | 1 |
| Last Updated | 2026-03-16 |
Key capabilities associated with this model.
Representative applications attached to this model record.
Balanced assessment surfaced in static HTML.
Primary references preserved in static HTML for citation extraction.
| Reference | Link |
|---|---|
| Zhen et al., 2024. 3D-VLA: A 3D Vision-Language-Action Generative World Model. | Open source |
Nearby models linked from the current editorial record.
| Model | Category | World Model Type | Index v1.1 |
|---|---|---|---|
| AMI World Model | Foundation World Model | Multimodal generative world model | 38/100 |
| TD-MPC2 | Model-Based RL | Implicit dynamics + MPC planner | 80/100 |
| UniSim | Generative World Model | Action-conditioned video simulator | 72/100 |
Side-by-side comparisons already connected to this model.
| Comparison | Matchup | Summary |
|---|---|---|
| 3D-VLA vs I-JEPA | 3D-VLA vs I-JEPA | Two approaches to learning representations for embodied intelligence: 3D-VLA combines 3D perception with language-conditioned action planning, while I-JEPA learns abstract visual representations through self-supervised prediction in latent space. |
| RT-2 vs 3D-VLA | RT-2 vs 3D-VLA | Two approaches to vision-language-action models for robotics. RT-2 leverages web-scale VLM knowledge through action tokenization, while 3D-VLA integrates explicit 3D spatial understanding for embodied reasoning. |
Connected research areas surfaced directly in static HTML.
| Topic | Summary |
|---|---|
| Language-Conditioned World Models | How language-conditioned world models use text prompts or natural-language actions to control simulation, planning, and embodied behavior across Pandora, 3D-VLA, RT-2, and hybrid systems. |
FAQ answers rendered directly into static HTML for extractable responses.
The world model enables the robot to predict consequences of actions before executing them, improving safety and planning capability in real-world tasks.
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.
This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-03-16.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary model and lab sources embedded in static HTML.