Static HTML snapshot of the model record for crawlers and no-JS readers.
| Attribute | Value |
|---|---|
| Model | STEVE-1 |
| Lab / Organization | UT Austin |
| Category | Generative World Model |
| Subtype | Instruction-Following Game Agent |
| World Model Type | Text-conditioned generative agent in open-world |
| Primary Domain | Game AI / Open World |
| Architecture | VPT backbone + MineCLIP goal encoder + latent-conditioned policy |
| Modality | Visual + Language (instructions) |
| Training Method | Hindsight relabeling on unlabeled gameplay + CLIP-based instruction conditioning |
| Status | active |
| Year | 2023 |
| Performance Index | 53/100 (medium confidence, v1.1) |
Main editorial body preserved directly in static HTML.
STEVE-1 is a generative agent for Minecraft that follows open-ended text and visual instructions. Built on top of a VPT (Video Pre-Training) backbone, STEVE-1 uses a latent-conditioned policy that translates CLIP-based goal embeddings into action sequences. It demonstrates that combining large-scale video pre-training with instruction-conditioned generation creates agents capable of executing diverse, compositional tasks in open-world environments.
STEVE-1 is a text-conditioned generative agent in open-world developed by UT Austin in 2023 for game ai / open world.
Short extractable facts for answer engines and no-JS readers.
| Signal | Value |
|---|---|
| Definition | STEVE-1 is a text-conditioned generative agent in open-world developed by UT Austin in 2023 for game ai / open world. |
| Short Description | An instruction-following agent for Minecraft that uses a generative world model to execute open-ended text commands. |
| Benchmark Rows | 1 |
| FAQ Entries | 2 |
| Related Models | 4 |
| Related Guides | 0 |
| Related Research Topics | 0 |
| Last Updated | 2026-04-07 |
Key capabilities associated with this model.
Representative applications attached to this model record.
Balanced assessment surfaced in static HTML.
Published benchmark evidence attached to this model record.
| Benchmark | Metric | Result | Source |
|---|---|---|---|
| Programmatic Eval | Task Completion | Strong instruction-following | Source |
Primary references preserved in static HTML for citation extraction.
| Reference | Link |
|---|---|
| Lifshitz et al., 2023. STEVE-1: A Generative Model for Text-to-Behavior in Minecraft. arXiv:2306.00937 | Open source |
Nearby models linked from the current editorial record.
| Model | Category | World Model Type | Index v1.1 |
|---|---|---|---|
| DreamerV3 | Model-Based RL | Imagination-based dynamics model | 88/100 |
| Genie | Generative World Model | Action-controllable generative model | 57/100 |
| OASIS | Generative World Model | Real-time playable world model | 66/100 |
| GameNGen | Generative World Model | Diffusion-based neural game engine | 52/100 |
Side-by-side comparisons already connected to this model.
| Comparison | Matchup | Summary |
|---|---|---|
| STEVE-1 vs DreamerV3 | STEVE-1 vs DreamerV3 | Two approaches to open-world game AI. STEVE-1 uses video pre-training and instruction following, while DreamerV3 learns a world model from scratch via reinforcement learning. |
FAQ answers rendered directly into static HTML for extractable responses.
DreamerV3 learns from scratch via RL and a world model. STEVE-1 uses large-scale video pre-training (VPT) and instruction conditioning, requiring no reward function but needing more pre-training data.
The architecture is Minecraft-specific, but the approach of combining video pre-training with instruction conditioning is applicable to any visual domain with abundant video data.
Short extractable summary preserved directly in static HTML.
Editorial provenance and refresh policy preserved directly in static HTML.
Published by world-models.io editorial board.
Lead editor Tyler D. - Technical editor, methodology and benchmark analysis.
This model page synthesizes primary papers, official model pages, benchmark evidence, and related world-models.io context into a reference resource.
Each editorial page is assembled from primary sources, normalized into extractable summaries, checked for factual drift, and reviewed before publication or major refreshes. Last reviewed: 2026-04-07.
Pages are refreshed when a new paper, benchmark, release, architecture update, or stronger primary source materially changes the answer a reader or AI system should retrieve.
Each page links back to relevant primary sources and keeps a stable canonical URL so readers can verify claims, trace context, and reference the most up-to-date version. See the editorial policy.
Primary model and lab sources embedded in static HTML.