Abstract
World models are becoming one of the most interesting and one of the most confusing ideas in artificial intelligence.
At a simple level, a world model is an AI system that tries to learn how an environment actually works. It can predict what may happen next, simulate possible futures, support planning, or help an agent understand the consequences of its actions. In practice, the field covers reinforcement learning, robotics, autonomous driving, generative video, embodied AI, simulation, spatial intelligence and autonomous agents.
This report provides a structured overview of the world models landscape as it stands in 2026. It does not claim that one definition is the only correct one. Instead, it organizes the field, explains the main model families, compares emerging benchmarks and identifies open problems.
The goal is to provide a practical reference for people who want to understand the field without getting lost in hype, jargon, or overly narrow technical definitions.
HTML edition
A practical taxonomy
The world models landscape can be organized around several dimensions. None is sufficient on its own, but together they provide a workable framework for comparison.
The five axes are domain, function, representation, time horizon and action conditioning. The fifth is especially important for robotics, autonomous driving and agent planning because a model that cannot represent the effect of actions cannot reliably help an agent act.
Six model families
The report distinguishes reinforcement learning, video, embodied, autonomous driving, spatial and 3D, and agentic or procedural world models.
These families serve different purposes. A visually rich video model may remain weak for planning, while a less visual reinforcement learning model may be more useful inside a decision loop. Universal rankings should therefore be treated with caution.
Benchmarks and evaluation
World models are difficult to benchmark because no single task covers the entire field. WorldModelBench focuses on video generation as world modeling, WorldPrediction on high-level world modeling and long-horizon procedural planning, and WorldArena on the gap between perceptual quality and functional utility in embodied systems.
Evaluation should include temporal coherence, physical consistency, object permanence, action sensitivity, causal plausibility, planning utility, generalization and functional value. A model should not only be beautiful; it should be useful.
Open challenges
Seven recurring problems remain unresolved: unclear definitions, the mistaken equation of video generation with world modeling, weak long-horizon reasoning, underdeveloped action conditioning, difficult sim-to-real transfer, scarce multimodal interaction data, and safety under uncertain predictions.
The broader signal is convergence. Reinforcement learning, robotics, video generation, autonomous driving, simulation and agents are moving toward the same question: can AI systems learn useful predictive models of the environments they operate in?
Conclusion
World models represent a shift from generating outputs toward learning environments, predicting change, simulating futures and supporting action.
Their future will not be measured only by how real they look, but by how well they help systems understand, plan and act.
Methodology
This report is based on a structured review of public research papers, benchmark pages, model documentation and technical project pages.
The objective is not to provide a complete academic survey. It is to provide a practical map of the main ideas, families, benchmarks and open problems in a field that is both broad and moving quickly.
The field is moving quickly, some systems are proprietary, some claims are difficult to verify, benchmarks are young and the term world model is used differently across communities. The report is a structured snapshot, not a final authority.
Bibliography
- Ha, D., & Schmidhuber, J. (2018). World Models. arXiv:1803.10122.
- Li, D., et al. (2025). WorldModelBench: Judging Video Generation Models As World Models. arXiv:2502.20694.
- Chen, D., Chung, W., Bang, Y., Ji, Z., & Fung, P. (2025). WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning. arXiv:2506.04363.
- Shang, Y., et al. (2026). WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models. arXiv:2602.08971.
- WorldModelBench project page. Judging Video Generation Models As World Models. See reference 02.
- WorldPrediction project page. A Benchmark for High-Level World Modeling and Long-Horizon Procedural Planning. See reference 03.
- WorldArena project page. A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models. See reference 04.
- WorldArena 2.0 project page. Extending Embodied World Model Benchmarking on Modality, Functionality and Platform.
Cite this report
GRENAT, Bernard. (2026). State of World Models 2026: Taxonomy, Benchmarks and Open Challenges (Version 1.0). world-models.io. Zenodo. https://doi.org/10.5281/zenodo.21345187
Version & changelog
Author: Bernard GRENAT. DOI: 10.5281/zenodo.21345187. License: CC BY 4.0.