NTH

Dream-Cubed: Controllable Generative Modeling in Minecraft by Training on Billions of Cubes

AuthorsTim Merino, Sam Earle, Ryunosuke Iwai, Julian Togelius, Edoardo Cetin

June 9, 2026 2 min read
Watch on YouTube
The one-line take

Dream-Cubed turns Minecraft worlds into a massive training ground for learning controllable 3D generative models that can fill in, extend, and create block-based environments interactively.

Key results

2026543
chunks

Total Dream-Cubed dataset size across procedural and human-authored Minecraft chunks

15
biomes

Number of biome labels in the core natural dataset

280M
parameters

Shared 3D Diffusion Transformer backbone size

59.26
MD4 FID

Adjusted average render-based FID for discrete masked diffusion

59.29
DDPM FID

Adjusted average render-based FID for continuous diffusion

19
participants

Minecraft-experienced participants in the human preference study

What the paper found

Dream-Cubed, from New York University and Sakana AI, is a block-level generative modeling paper that turns Minecraft into a controllable 3D diffusion benchmark by training on tens of billions of cubes. The core contribution is a curated dataset of 2,026,543 chunks at 32×32×32 resolution, combining 1,667,781 procedurally generated terrain chunks across 15 biome labels with 358,762 human-authored chunks from six professional maps, then using that corpus to compare a 280M-parameter 3D Diffusion Transformer under two formulations: discrete masked diffusion (MD4) and continuous DDPM in embedding space. The models perform almost identically on adjusted render-based FID, 59.26 for MD4 and 59.29 for DDPM, but the discrete formulation uniquely enables hard-constraint inpainting, outpainting, and user-authored block prompting without architectural changes. A human preference study with 19 Minecraft-experienced participants found generated chunks were preferred over real validation chunks for MD4 patch 2 at 67.1%, MD4 patch 4 at 57.1%, and DDPM patch 2 at 55.2%, while MD4 patch 2 narrowly beat DDPM patch 2 in direct comparison at 49.4%, showing that patch size matters more than diffusion type. The paper also shows that sampling quality degrades sharply for MD4 when reducing steps from 1,000 to 10, with average FID rising from 59.26 to 174.53, and that targeted data composition improves structured biomes such as villages and caves relative to naive frequency-based sampling.

Original abstract

We introduce Dream-Cubed, a large-scale dataset of Minecraft worlds at voxel resolution, and a family of models using cubes as powerful compositional units for efficient generation of interactive 3D environments. Dream-Cubed comprises tens of billions of tokens from a carefully curated mixture of procedural biome terrain and high-quality human-authored maps. We use this dataset to conduct the first large-scale study of 3D diffusion models for voxel generation, analyzing discrete and continuous diffusion formulations, data compositions, and architectural design choices. Our models operate directly in the space of blocks, enabling efficient and semantically grounded generation while supporting interactive user workflows such as inpainting and outpainting from user-authored blocks. To quantitatively evaluate our models, we adapt the FID metric to assess semantic differences between real and generated world renderings, and validate generation quality through a human preference study. We release the full dataset, code, and all our pretrained models, which we hope will provide a foundation for future research in efficient generative modeling for structured, interactive 3D environments.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis