New: the Timeline is live. Track world model releases, papers, and benchmark updates in real time.
world-models.io
The Knowledge Hub for AI World Models

STEVE-1

STEVE-1 is a generative agent for Minecraft that follows text and visual instructions, using a latent world model to navigate and interact in open-ended 3D environments.

robotics model-based-rl simulation embodied-ai

Key Attributes

Static HTML snapshot of the model record for crawlers and no-JS readers.

AttributeValue
ModelSTEVE-1
Lab / OrganizationUT Austin
CategoryGenerative World Model
SubtypeInstruction-Following Game Agent
World Model TypeText-conditioned generative agent in open-world
Primary DomainGame AI / Open World
ArchitectureVPT backbone + MineCLIP goal encoder + latent-conditioned policy
ModalityVisual + Language (instructions)
Training MethodHindsight relabeling on unlabeled gameplay + CLIP-based instruction conditioning
Statusactive
Year2023
Performance Index53/100 (medium confidence, v1.1)

About STEVE-1

Main editorial body preserved directly in static HTML.

STEVE-1 is a generative agent for Minecraft that follows open-ended text and visual instructions. Built on top of a VPT (Video Pre-Training) backbone, STEVE-1 uses a latent-conditioned policy that translates CLIP-based goal embeddings into action sequences. It demonstrates that combining large-scale video pre-training with instruction-conditioned generation creates agents capable of executing diverse, compositional tasks in open-world environments.

STEVE-1 is a text-conditioned generative agent in open-world developed by UT Austin in 2023 for game ai / open world.

Editorial Snapshot

Short extractable facts for answer engines and no-JS readers.

SignalValue
DefinitionSTEVE-1 is a text-conditioned generative agent in open-world developed by UT Austin in 2023 for game ai / open world.
Short DescriptionAn instruction-following agent for Minecraft that uses a generative world model to execute open-ended text commands.
Benchmark Rows1
FAQ Entries2
Related Models4
Related Guides0
Related Research Topics0
Last Updated2026-04-07

Notable Features

Key capabilities associated with this model.

  • Open-ended instruction following
  • Minecraft open-world navigation
  • CLIP-based goal conditioning
  • No reward function needed

Use Cases

Representative applications attached to this model record.

Open-world game playingInstruction-following agentsCreative task executionEmbodied language grounding

Strengths and Limitations

Balanced assessment surfaced in static HTML.

Strengths

  • Handles diverse open-ended instructions
  • No reward engineering needed
  • Scales with video pre-training data
  • Compositional task understanding

Limitations

  • Minecraft-specific
  • Limited long-horizon planning
  • Requires large pre-training corpus
  • Short action horizons

Benchmarks

Published benchmark evidence attached to this model record.

BenchmarkMetricResultSource
Programmatic EvalTask CompletionStrong instruction-followingSource

References and Citations

Primary references preserved in static HTML for citation extraction.

ReferenceLink
Lifshitz et al., 2023. STEVE-1: A Generative Model for Text-to-Behavior in Minecraft. arXiv:2306.00937Open source

Related Models

Nearby models linked from the current editorial record.

ModelCategoryWorld Model TypeIndex v1.1
DreamerV3Model-Based RLImagination-based dynamics model88/100
GenieGenerative World ModelAction-controllable generative model57/100
OASISGenerative World ModelReal-time playable world model66/100
GameNGenGenerative World ModelDiffusion-based neural game engine52/100

Direct Comparisons

Side-by-side comparisons already connected to this model.

ComparisonMatchupSummary
STEVE-1 vs DreamerV3STEVE-1 vs DreamerV3Two approaches to open-world game AI. STEVE-1 uses video pre-training and instruction following, while DreamerV3 learns a world model from scratch via reinforcement learning.

Frequently Asked Questions

FAQ answers rendered directly into static HTML for extractable responses.

How is STEVE-1 different from DreamerV3 in Minecraft?

DreamerV3 learns from scratch via RL and a world model. STEVE-1 uses large-scale video pre-training (VPT) and instruction conditioning, requiring no reward function but needing more pre-training data.

Can STEVE-1 be applied outside Minecraft?

The architecture is Minecraft-specific, but the approach of combining video pre-training with instruction conditioning is applicable to any visual domain with abundant video data.

Quick Answer

Short extractable summary preserved directly in static HTML.

  • STEVE-1 is a text-conditioned generative agent in open-world developed by UT Austin in 2023 for game ai / open world.
  • Use this page when you need a fast read on how STEVE-1 fits into the generative world model landscape, then validate the details in the benchmarks, citations, and related pages.
  • A key strength surfaced in the editorial record is handles diverse open-ended instructions.

Editorial Trust Signals

Editorial provenance and refresh policy preserved directly in static HTML.

Published by world-models.io editorial board.

Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.

This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.

Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-04-07.

Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.

Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.

Primary sources onlyLast reviewed date visibleMethodology documentedSource links included

External Sources

Primary model and lab sources embedded in static HTML.

References

  1. [1] Lifshitz et al., 2023. STEVE-1: A Generative Model for Text-to-Behavior in Minecraft. arXiv:2306.00937