NTH

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation

AuthorsAlexi Gladstone, Heng Ji, Yilun Du

August 1, 2026 3 min read
Watch on YouTube
The one-line take

Explorative Modeling trains generative systems by trying multiple candidate outputs and learning from the best match, promising better scaling and much faster end-to-end generation.

Key results

4.1×
FLOP efficiency improvement

Exploration reaches baseline image-generation performance with 4.1× fewer FLOPs.

6.2×
Sample efficiency improvement

The RAE ImageNet recipe reaches baseline performance with 6.2× less training data.

47%
Parameter efficiency improvement

A Large model with exploration outscales an XLarge model with 47% more parameters and no exploration.

1.43
Unguided ImageNet FID

XRAE achieves a near-state-of-the-art unguided FID on ImageNet.

256×
Inference reduction

Explorative World Models match Diffuser on Maze2D with up to 256× fewer inference evaluations.

What the paper found

Researchers Alexi Gladstone and Heng Ji at the University of Illinois Urbana-Champaign, with Yilun Du at Harvard, introduce Explorative Modeling, or XM, as a third generative-model scaling axis alongside parameters and data. Instead of factoring generation into many autoregressive or diffusion steps, XM factors the training loop: it samples K candidate outputs, matches each to real data, and backpropagates only through the lowest-loss candidate. Forward XM improves coverage, while Reverse XM improves precision but requires entropy or coverage control to avoid collapse. Added to Diffusion, Flow Matching, Jumpy models, and masked diffusion language models, exploration improves image, video, and language generation, with gains increasing at scale from 7% to 36% as data grows and from 13% to 23% as model size grows. On the RAE ImageNet recipe, exploration delivers 4.1× better FLOP efficiency, 6.2× better sample efficiency, 47% better parameter efficiency, and an unguided 1.43 FID. The approach also makes reconstructive generation genuinely end-to-end: an Explorative Policy matches Diffusion Policy on Robomimic with one forward pass instead of 100, while an Explorative World Model matches Diffuser on Maze2D using up to 256× fewer inference evaluations. Unlike conventional staged generation, including systems such as OpenAI’s GPT-4-style autoregressive modeling and diffusion pipelines, XM shifts computation from inference to training, reducing exposure bias while expanding the number of modes a model can represent.

Original abstract

The deep learning revolution, kicked off by AlexNet, taught us that end-to-end training beats decomposing a problem into hand-designed stages. Generative modeling, however, has remained the exception-despite generative models being remarkably capable, they are still not trained end-to-end. This is because, at its core, generative modeling is about handling distributions with many modes, and existing scalable approaches handle this the same way, by factoring the generation procedure, which prevents end-to-end generation. In this work, we introduce Explorative Modeling, a new paradigm that instead factors the training loop, exploring K candidate matches between model generations and data, and training on the best, so predictions commit to modes rather than blurring them. We find Explorative Models (XMs) useful in two settings. First, increasing exploration adds a third pretraining axis beyond parameters and data for existing generative models-where scaling exploration monotonically improves performance across both continuous and discrete domains (images, video, and language). Notably, gains from exploration increase with scale, climbing from 7% to 36% as data scales and from 13% to 23% as models grow, with efficiency gains more than doubling at 3x the compute. Concretely, exploration improves FLOP efficiency by 4.1x, sample efficiency by 6.2x, parameter efficiency by 47%, lifts the strongest of image-generation recipes to a near-state-of-the-art 1.43 FID on ImageNet without guidance, enables scaling how end-to-end existing models are, and unlocks scaling generalization. Second, XMs enable end-to-end reconstructive generative modeling, matching diffusion on control tasks with 16-256x fewer inference steps. Together, these results establish XMs as both a new pretraining axis for existing generative models and a standalone end-to-end generative modeling paradigm.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis