Static HTML snapshot of the model record for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Model | RT-2 |
| Lab / Organization | DeepMind |
| Category | Foundation World Model |
| Subtype | Vision-Language-Action Model |
| World Model Type | Web-knowledge transfer model for robotics |
| Primary Domain | Robotics |
| Architecture | PaLI-X / PaLM-E backbone fine-tuned with action tokenization |
| Modality | Visual + Language + Proprioceptive |
| Training Method | Co-fine-tuning on web VLM data + robot demonstration trajectories |
| Status | active |
| Year | 2023 |
| Performance Index | 72/100 (medium confidence, v1.1) |
Main editorial body preserved directly in static HTML.
RT-2 (Robotic Transformer 2) is a vision-language-action (VLA) model that co-fine-tunes a large vision-language model on both web data and robotic trajectories. By tokenizing robot actions as text tokens, RT-2 enables a single model to perform visual question answering, image captioning, and robot control simultaneously. It demonstrates emergent capabilities such as reasoning about previously unseen objects, interpreting abstract instructions, and performing multi-step manipulation tasks, all without task-specific training.
RT-2 is a web-knowledge transfer model for robotics developed by Google DeepMind in 2023 for robotics.
Short extractable facts for answer engines and no-JS readers.
| Signal | Value |
|---|---|
| Definition | RT-2 is a web-knowledge transfer model for robotics developed by Google DeepMind in 2023 for robotics. |
| Short Description | A vision-language-action model that transfers web-scale knowledge directly to robot control. |
| Benchmark Rows | 2 |
| FAQ Entries | 2 |
| Related Models | 4 |
| Related Guides | 0 |
| Related Research Topics | 1 |
| Last Updated | 2026-04-07 |
Key capabilities associated with this model.
Representative applications attached to this model record.
Balanced assessment surfaced in static HTML.
Primary references preserved in static HTML for citation extraction.
| Reference | Link |
|---|---|
| Brohan et al., 2023. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. arXiv:2307.15818 | Open source |
Nearby models linked from the current editorial record.
| Model | Category | World Model Type | Index v1.1 |
|---|---|---|---|
| 3D-VLA | Foundation World Model | 3D-aware embodied world model | 47/100 |
| AMI World Model | Foundation World Model | Multimodal generative world model | 38/100 |
| TD-MPC2 | Model-Based RL | Implicit dynamics + MPC planner | 80/100 |
Side-by-side comparisons already connected to this model.
| Comparison | Matchup | Summary |
|---|---|---|
| RT-2 vs 3D-VLA | RT-2 vs 3D-VLA | Two approaches to vision-language-action models for robotics. RT-2 leverages web-scale VLM knowledge through action tokenization, while 3D-VLA integrates explicit 3D spatial understanding for embodied reasoning. |
Connected research areas surfaced directly in static HTML.
| Topic | Summary |
|---|---|
| Language-Conditioned World Models | How language-conditioned world models use text prompts or natural-language actions to control simulation, planning, and embodied behavior across Pandora, 3D-VLA, RT-2, and hybrid systems. |
FAQ answers rendered directly into static HTML for extractable responses.
RT-1 was trained only on robot data. RT-2 co-fine-tunes a large vision-language model (PaLI-X or PaLM-E) on both web and robot data, enabling emergent reasoning capabilities that RT-1 lacks.
Yes. By leveraging web-scale visual knowledge, RT-2 can reason about and manipulate objects not present in its robot training data, a key emergent capability.
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.
This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-04-07.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary model and lab sources embedded in static HTML.