New: the Timeline is live. Track world model releases, papers, and benchmark updates in real time.
world-models.io
The Knowledge Hub for AI World Models

TD-MPC2

TD-MPC2 combines temporal-difference learning with model-predictive control to create a scalable world model agent capable of mastering 104+ continuous control tasks.

robotics model-based-rl simulation embodied-ai

Key Attributes

Static HTML snapshot of the model record for crawlers and no-JS readers.

AttributeValue
ModelTD-MPC2
Lab / OrganizationMIT
CategoryModel-Based RL
SubtypeTemporal Difference Learning + MPC
World Model TypeImplicit dynamics + MPC planner
Primary DomainMulti-task control
ArchitectureImplicit latent dynamics model with TD-learning and MPPI planning
ModalityState + Visual
Training MethodJoint TD-learning and latent dynamics model training with MPC-based acting
Statusactive
Year2024
Performance Index80/100 (high confidence, v1.1)

About TD-MPC2

Main editorial body preserved directly in static HTML.

TD-MPC2 scales model-based RL to a single generalist agent that masters 104 tasks across multiple domains, combining temporal difference learning with model-predictive control. It demonstrates that a single world model can generalize across vastly different task types while maintaining strong performance.

TD-MPC2 is an implicit dynamics + mpc planner developed by MIT / Meta in 2024 for multi-task control.

Editorial Snapshot

Short extractable facts for answer engines and no-JS readers.

SignalValue
DefinitionTD-MPC2 is an implicit dynamics + mpc planner developed by MIT / Meta in 2024 for multi-task control.
Short DescriptionA scalable world model agent that combines TD-learning with model-predictive control across 104 diverse tasks.
Benchmark Rows2
FAQ Entries1
Related Models3
Related Guides3
Related Research Topics3
Last Updated2026-03-05

Notable Features

Key capabilities associated with this model.

  • Single agent for 104 diverse tasks
  • Combines TD-learning with MPC
  • Task-conditioned predictions
  • Strong multi-task generalization

Use Cases

Representative applications attached to this model record.

Multi-task controlRoboticsLocomotionManipulation

Strengths and Limitations

Balanced assessment surfaced in static HTML.

Strengths

  • Multi-task generalization
  • Scales to 104 tasks
  • Strong sample efficiency
  • Combines learning and planning

Limitations

  • MPC planning cost at inference
  • Requires diverse training data
  • Complex training pipeline

Benchmarks

Published benchmark evidence attached to this model record.

BenchmarkMetricResultSource
DMControl (30 tasks)Mean Return 879 avg returnState-of-the-artSource
Meta-World (50 tasks)Success Rate 82 %Strong generalizationSource

References and Citations

Primary references preserved in static HTML for citation extraction.

ReferenceLink
Hansen et al., 2024. TD-MPC2: Scalable, Robust World Models for Continuous Control. ICLR 2024.Open source

Related Models

Nearby models linked from the current editorial record.

ModelCategoryWorld Model TypeIndex v1.1
DreamerV3Model-Based RLImagination-based dynamics model88/100
MuZeroModel-Based RLAbstract learned dynamics + MCTS78/100
PlaNetModel-Based RLLatent space planning model57/100

Direct Comparisons

Side-by-side comparisons already connected to this model.

ComparisonMatchupSummary
Model-Based RL vs Model-Free RLModel-Based RL vs Model-Free RLModel-based RL learns a world model for imagination-based planning. Model-free RL learns directly from interaction without an internal model. Each approach has distinct strengths depending on the application domain.
DreamerV3 vs TD-MPC2DreamerV3 vs TD-MPC2Two leading model-based RL agents with different philosophies: DreamerV3 uses imagination-based actor-critic learning, while TD-MPC2 combines temporal-difference learning with model-predictive control for multi-task mastery.
MuZero vs TD-MPC2MuZero vs TD-MPC2Both use learned dynamics models for planning, but MuZero uses Monte Carlo tree search for deep discrete planning while TD-MPC2 uses model-predictive control for continuous multi-task settings.
MuZero vs DreamerV3MuZero vs DreamerV3Two titans of model-based RL with fundamentally different approaches: MuZero learns a value-equivalent model for search-based planning, while DreamerV3 learns a generative world model for imagination-based policy optimization.
RT-2 vs 3D-VLART-2 vs 3D-VLATwo approaches to vision-language-action models for robotics. RT-2 leverages web-scale VLM knowledge through action tokenization, while 3D-VLA integrates explicit 3D spatial understanding for embodied reasoning.
LeWorldModel vs DreamerV3LeWorldModel vs DreamerV3LeWorldModel revisits LeCun's energy-based JEPA philosophy for control, predicting in latent space without pixel reconstruction. DreamerV3 remains the canonical RSSM-based agent that learns by imagining pixel-grounded rollouts.
PlayWorld vs TD-MPC2PlayWorld vs TD-MPC2Two green-index models for robot decision-making, but with very different operating modes. PlayWorld learns a manipulation-focused world simulator from autonomous play, while TD-MPC2 combines latent dynamics with model-predictive control across a wide multi-task control benchmark suite.
PlayWorld vs V-JEPA 2PlayWorld vs V-JEPA 2Two green-index models pushing robotics-relevant world understanding in different ways. PlayWorld is a robot-play simulator for manipulation and policy improvement, while V-JEPA 2 is a self-supervised video predictor optimized for physical reasoning and zero-shot robot planning.

Guides Referencing This Model

Crawler-readable guide links tied to this model.

GuideSummary
Building World Models: A Practical GuideA practical guide to implementing world models: from choosing architectures and training setups to debugging dynamics learning and policy optimization.
World Models for RoboticsHow to use world models for robot learning: from simulation-based training to real-world deployment and sim-to-real transfer.
Evaluating World Models: Benchmarks, Metrics and PitfallsA practical reference on how to benchmark world models: from sample efficiency on Atari 100K and DMControl to long-horizon prediction quality and downstream policy performance.

Research Topics Referencing This Model

Connected research areas surfaced directly in static HTML.

TopicSummary
Model-Based Reinforcement LearningWhat model-based reinforcement learning is, how world models enable imagination-based planning, and why Dreamer, MuZero, PlaNet, and TD-MPC2 matter.
World Models for RoboticsHow world models improve robot learning, learned simulation, safe exploration, and sim-to-real transfer across manipulation, navigation, and control.
World Model EvaluationHow to evaluate world models across rollout quality, benchmark performance, planning utility, and downstream transfer instead of relying on visual plausibility alone.

Frequently Asked Questions

FAQ answers rendered directly into static HTML for extractable responses.

How does TD-MPC2 scale to 104 tasks?

It uses a single latent dynamics model shared across all tasks, combined with task-conditioned predictions and MPPI planning.

Quick Answer

Short extractable summary preserved directly in static HTML.

  • TD-MPC2 is an implicit dynamics + mpc planner developed by MIT / Meta in 2024 for multi-task control.
  • Use this page when you need a fast read on how TD-MPC2 fits into the model-based rl landscape, then validate the details in the benchmarks, citations, and related pages.
  • A key strength surfaced in the editorial record is multi-task generalization.

Editorial Trust Signals

Editorial provenance and refresh policy preserved directly in static HTML.

Published by world-models.io editorial board.

Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.

This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.

Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-03-05.

Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.

Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.

Primary sources onlyLast reviewed date visibleMethodology documentedSource links included

External Sources

Primary model and lab sources embedded in static HTML.

References

  1. [1] Hansen et al., 2024. TD-MPC2: Scalable, Robust World Models for Continuous Control. ICLR 2024.