New: the Timeline is live. Track world model releases, papers, and benchmark updates in real time.
world-models.io
The Knowledge Hub for AI World Models

3D-VLA

3D-VLA integrates 3D vision, language understanding, and action prediction into a unified model for embodied AI tasks.

robotics model-based-rl simulation embodied-ai

Key Attributes

Static HTML snapshot of the model record for crawlers and no-JS readers.

AttributeValue
Model3D-VLA
Lab / OrganizationMIT
CategoryFoundation World Model
Subtype3D Vision-Language-Action Model
World Model Type3D-aware embodied world model
Primary DomainRobotics / Embodied AI
Architecture3D-aware vision-language model with integrated world model and action decoder
Modality3D Point Clouds + Language + Actions
Training MethodMulti-task training: 3D understanding, world modeling, and action generation
Statusemerging
Year2024
Performance Index47/100 (low confidence, v1.1)

About 3D-VLA

Main editorial body preserved directly in static HTML.

3D-VLA integrates 3D perception, language understanding, and action generation with an explicit world model component. The model can reason about 3D scenes, predict future states, and generate robot actions, all within a unified architecture. By incorporating a world model, 3D-VLA can plan ahead by imagining consequences of actions in 3D space before executing them.

3D-VLA is a 3d-aware embodied world model developed by MIT / Tsinghua in 2024 for robotics / embodied ai.

Editorial Snapshot

Short extractable facts for answer engines and no-JS readers.

SignalValue
Definition3D-VLA is a 3d-aware embodied world model developed by MIT / Tsinghua in 2024 for robotics / embodied ai.
Short DescriptionA 3D vision-language-action model with a built-in world model for embodied AI, enabling 3D-aware reasoning, planning, and action generation.
Benchmark Rows0
FAQ Entries1
Related Models3
Related Guides0
Related Research Topics1
Last Updated2026-03-16

Notable Features

Key capabilities associated with this model.

  • Unified 3D perception + world model + action
  • 3D-aware future prediction
  • Language-conditioned planning
  • Embodied task execution

Use Cases

Representative applications attached to this model record.

Robot manipulation3D scene understandingLanguage-guided roboticsEmbodied planning

Strengths and Limitations

Balanced assessment surfaced in static HTML.

Strengths

  • Unified architecture
  • 3D-native understanding
  • Integrated world model for planning
  • Language-conditioned

Limitations

  • Early-stage research
  • Limited real-robot validation
  • High compute requirements

References and Citations

Primary references preserved in static HTML for citation extraction.

ReferenceLink
Zhen et al., 2024. 3D-VLA: A 3D Vision-Language-Action Generative World Model.Open source

Related Models

Nearby models linked from the current editorial record.

ModelCategoryWorld Model TypeIndex v1.1
AMI World ModelFoundation World ModelMultimodal generative world model38/100
TD-MPC2Model-Based RLImplicit dynamics + MPC planner80/100
UniSimGenerative World ModelAction-conditioned video simulator72/100

Direct Comparisons

Side-by-side comparisons already connected to this model.

ComparisonMatchupSummary
3D-VLA vs I-JEPA3D-VLA vs I-JEPATwo approaches to learning representations for embodied intelligence: 3D-VLA combines 3D perception with language-conditioned action planning, while I-JEPA learns abstract visual representations through self-supervised prediction in latent space.
RT-2 vs 3D-VLART-2 vs 3D-VLATwo approaches to vision-language-action models for robotics. RT-2 leverages web-scale VLM knowledge through action tokenization, while 3D-VLA integrates explicit 3D spatial understanding for embodied reasoning.

Research Topics Referencing This Model

Connected research areas surfaced directly in static HTML.

TopicSummary
Language-Conditioned World ModelsHow language-conditioned world models use text prompts or natural-language actions to control simulation, planning, and embodied behavior across Pandora, 3D-VLA, RT-2, and hybrid systems.

Frequently Asked Questions

FAQ answers rendered directly into static HTML for extractable responses.

Why integrate a world model into a VLA?

The world model enables the robot to predict consequences of actions before executing them, improving safety and planning capability in real-world tasks.

Quick Answer

Short extractable summary preserved directly in static HTML.

  • 3D-VLA is a 3d-aware embodied world model developed by MIT / Tsinghua in 2024 for robotics / embodied ai.
  • Use this page when you need a fast read on how 3D-VLA fits into the foundation world model landscape, then validate the details in the benchmarks, citations, and related pages.
  • A key strength surfaced in the editorial record is unified architecture.

Editorial Trust Signals

Editorial provenance and refresh policy preserved directly in static HTML.

Published by world-models.io editorial board.

Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.

This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.

Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-03-16.

Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.

Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.

Primary sources onlyLast reviewed date visibleMethodology documentedSource links included

External Sources

Primary model and lab sources embedded in static HTML.

References

  1. [1] Zhen et al., 2024. 3D-VLA: A 3D Vision-Language-Action Generative World Model.