Static HTML snapshot of the model record for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Model | Predictron |
| Lab / Organization | DeepMind |
| Category | Model-Based RL |
| Subtype | Abstract World Model |
| World Model Type | Abstract internal dynamics model |
| Primary Domain | Value prediction |
| Architecture | Multi-step abstract model with λ-weighted returns |
| Modality | Abstract state representations |
| Training Method | End-to-end supervised learning with multi-step abstract predictions |
| Status | foundational |
| Year | 2017 |
| Performance Index | 43/100 (medium confidence, v1.1) |
Main editorial body preserved directly in static HTML.
The Predictron combines learning and planning by performing multiple steps of abstract lookahead within a neural network. It learns abstract internal dynamics optimized directly for value prediction, without requiring explicit environment reconstruction. It was one of the earliest demonstrations that world models could be learned end-to-end.
Predictron is an abstract internal dynamics model developed by Google DeepMind in 2017 for value prediction.
Short extractable facts for answer engines and no-JS readers.
| Signal | Value |
|---|---|
| Definition | Predictron is an abstract internal dynamics model developed by Google DeepMind in 2017 for value prediction. |
| Short Description | An architecture that integrates learning and planning into a single differentiable network via abstract world models. |
| Benchmark Rows | 1 |
| FAQ Entries | 1 |
| Related Models | 2 |
| Related Guides | 0 |
| Related Research Topics | 2 |
| Last Updated | 2026-01-28 |
Key capabilities associated with this model.
Representative applications attached to this model record.
Balanced assessment surfaced in static HTML.
Published benchmark evidence attached to this model record.
| Benchmark | Metric | Result | Source |
|---|---|---|---|
| Grid-world planning | RMSE 0.12 RMSE↓ | Strong improvement over model-free | Source |
Primary references preserved in static HTML for citation extraction.
| Reference | Link |
|---|---|
| Silver et al., 2017. The Predictron: End-to-End Learning and Planning. ICML 2017. | Open source |
Side-by-side comparisons already connected to this model.
| Comparison | Matchup | Summary |
|---|---|---|
| DreamerV3 vs MuZero | DreamerV3 vs MuZero | Both are landmark world model systems, but with fundamentally different architectures. DreamerV3 uses latent imagination with actor-critic learning, while MuZero uses abstract learned dynamics with Monte Carlo tree search. |
| MuZero vs TD-MPC2 | MuZero vs TD-MPC2 | Both use learned dynamics models for planning, but MuZero uses Monte Carlo tree search for deep discrete planning while TD-MPC2 uses model-predictive control for continuous multi-task settings. |
| Predictron vs MuZero | Predictron vs MuZero | Both learn abstract dynamics models for planning without requiring environment reconstruction, but Predictron was an early prototype while MuZero became the definitive realization of value-equivalent model learning. |
| Predictron vs MuZero | Predictron vs MuZero | Two DeepMind models that learn abstract value-equivalent dynamics. The Predictron (2017) introduced the concept of learned transition models in abstract space; MuZero (2020) scaled this to superhuman game play without knowing the rules. |
Connected research areas surfaced directly in static HTML.
| Topic | Summary |
|---|---|
| Model-Based Reinforcement Learning | What model-based reinforcement learning is, how world models enable imagination-based planning, and why Dreamer, MuZero, PlaNet, and TD-MPC2 matter. |
| World Models: A Comprehensive Survey | A survey of AI world models covering taxonomy, leading architectures, landmark systems, open challenges, and future research directions. |
FAQ answers rendered directly into static HTML for extractable responses.
Instead of learning an explicit environment model, the Predictron learns abstract internal dynamics optimized directly for value prediction, without needing to reconstruct observations.
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.
This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-01-28.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary model and lab sources embedded in static HTML.