Action-conditioned video world models for Minecraft. FlowCraft models are causal diffusion transformers trained with flow matching in the latent space of a fine-tuned WAN 2.2 VAE. They generate autoregressively, one latent at a time, conditioned on the latents and actions so far.
This organisation holds the data, the VAEs and the model checkpoints from How to Train Your Video World Model for Long-Horizon Rollout (arXiv link coming soon). The paper compares pre-training recipes at equal compute, then post-trains three of them (Diffusion Forcing, clean-prefix teacher forcing and chunked teacher forcing): flow map finetuning followed by Self Forcing against the model's own pre-trained teacher. Every stage is evaluated on 1000-frame rollouts, far beyond the attention window.
How to Train Your Video World Model for Long-Horizon Rollout
Aidan Scannell*, Samuel Garcin*, Paul Chang, Peter Bell and Amos Storkey
Toyota Research Institute · University of Edinburgh · Verda. *Equal contribution.
Code: github.com/aidanscannell/flowcraft
| Resolution | Latents | VAE | Models |
|---|---|---|---|
| 128×224 | minecraft-latents-128-224-v2 (142 GB) |
wan-vae-minecraft-128-224 |
all 16 checkpoints |
| 352×640 | minecraft-latents-352-640-v2 (1.1 TB) |
wan-vae-minecraft-352-640 |
none yet |
Collections: Models · VAEs · Datasets
The latents are the gameplay recordings from OpenAI's Video Pre-Training (VPT) contractor data, encoded once with the VAE of the same resolution, with the per-frame actions alongside. You never need the raw videos to train or evaluate.
{pre,post}-train-<recipe>[-quarter-budget]
post-train-df comes from pre-train-df, and so on).
Only stage 2 is released.pf, chunk, jump). -quarter-budget is 166 EFLOP
(85k / 45k steps).| Recipe | Paper name | What it trains on |
|---|---|---|
df |
Diffusion Forcing | An independent flow time for every latent, with a loss on each. Uniform flow-time prior. |
pf |
Clean-prefix teacher forcing | Every latent against its own clean prefix, all in one pass over a doubled clean/noisy sequence. |
chunk |
Chunked teacher forcing (B=4) | As pf, but targets come in blocks of 4 latents, each seeing the clean prefix and the earlier members of its block at their own flow times. |
jump |
Jump prediction (hmax=3) | As pf, but up to 3 latents between a target and its clean prefix are deleted instead of noised. |
ctx |
Context-length sampling | One context length per clip: the latents before it clean, a loss on every latent after it. |
ctx-notail |
ctx − tail | As ctx, with a loss only on the first latent after the cut. |
All recipes except df use a logit-normal flow-time prior. In pre-training,
every recipe except df also uses history corruption: on a fraction p=0.1 of clips the clean
history is presented at flow times τ ~ U[τmin, 1], with τmin=0.7.
A name without a suffix uses these defaults, so pre-train-pf and
pre-train-pf-quarter-budget differ only in budget. The suffixed variants
sweep the history corruption (quarter budget only):
| Suffix | History corruption |
|---|---|
pf-clean |
none (p=0) |
pf |
p=0.1, τmin=0.7 (default) |
pf-t05 |
p=0.1, τmin=0.5 |
pf-t00 |
p=0.1, τmin=0 |
Every checkpoint repo carries its full training config (config.json), the
step, and the sha256 of the original training checkpoint (meta.json).
git clone https://github.com/aidanscannell/flowcraft && cd flowcraft
uv sync --extra gpu --extra cu130
# Evaluate a released model (downloads the weights, VAE and test split)
uv run fc-eval ckpt_path=flow-craft/pre-train-pf
from omegaconf import OmegaConf
from flowcraft.config import ExperimentConfig
from flowcraft.model import build_world_model
from flowcraft.train.checkpoint import CheckpointManager, load_checkpoint_config
ref = "flow-craft/pre-train-pf@iclr2027-submission"
cfg = OmegaConf.merge(OmegaConf.structured(ExperimentConfig), load_checkpoint_config(ref))
model = build_world_model(cfg)
CheckpointManager(cfg.train.ckpt, "/tmp/unused").load(ref, model)
Only need a few tasks of the data?
from huggingface_hub import snapshot_download
snapshot_download(
"flow-craft/minecraft-latents-128-224-v2", repo_type="dataset",
allow_patterns=["build-house/**", "manifest.json", "splits/**"],
)
Every checkpoint has an iclr2027-submission tag: the exact version the
paper evaluated. Pin it (@iclr2027-submission) for reproduction; main may
move.
If you find FlowCraft, its checkpoints or its data useful, please consider citing us:
@article{scannell2026flowcraft,
title = {How to Train Your Video World Model for Long-Horizon Rollout},
author = {Scannell, Aidan and Garcin, Samuel and Chang, Paul and Bell, Peter and Storkey, Amos},
journal = {arXiv preprint arXiv:<ARXIV-ID>},
year = {2026}
}