World models evaluated through downstream control and sample efficiency on shared reinforcement-learning benchmarks.
Results require a documented task set, environment and ROM version, observation pipeline, seed count, interaction budget, Mean HNS, and human-normalized unit before they are comparable.
| Model | Metric | Result | Evidence |
|---|---|---|---|
| dreamer-v3 | Mean HNS | 2.01 x human | Author reported |
| iris | Mean HNS | 1.046 x human | Author reported |
| diamond | Mean HNS | 1.56 x human | Author reported |
Results require a documented DMControl task set, environment version, observation settings, seed count, and mean-return metric before they are comparable.
| Model | Metric | Result | Evidence |
|---|---|---|---|
| dreamer-v3 | Mean Return | 901 avg return | Author reported |
These source-linked author reports are retained for traceability but excluded from comparisons because they do not match a registered shared protocol.
| Model | Reported benchmark | Result | Why excluded |
|---|---|---|---|
| dreamer-v2 | Atari 200M | 1 x human | Source benchmark name does not match a registered public protocol. |
| planet | DMControl Suite | 10 x more efficient | Source benchmark name does not match a registered public protocol. |
| muzero | Atari | 7.31 x human | Source benchmark name does not match a registered public protocol. |
| td-mpc2 | DMControl (30 tasks) | 879 avg return | Source benchmark name does not match a registered public protocol. |
| imagination-augmented-agents | Atari | 1.15 x baseline | Source benchmark name does not match a registered public protocol. |
| value-prediction-network | Atari | 1.08 x baseline | Source benchmark name does not match a registered public protocol. |