NTH

Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation

AuthorsJiayi Xu, Di He, Guolin Ke

July 11, 2026 2 min read
Watch on YouTube
The one-line take

This paper makes pixel-by-pixel image generators train more like they infer, delivering a new state of the art without needing a separate tokenizer.

Key results

2.58
PRA-S FID

ImageNet-1K 256×256 class-conditional generation with 135M parameters

3.60
Previous best pixel-space AR FID

Prior billion-scale pixel-space autoregressive result on ImageNet-1K 256×256

1.94
PRA-L FID

ImageNet-1K 256×256 class-conditional generation with 511M parameters

16
Intermediate-state dimension

Default compact internal state size used by PRA

68.80
Linear probing top-1 accuracy

ImageNet-1K probing accuracy for PRA-L

What the paper found

Parallel Rollout Approximation, developed by Jiayi Xu, Di He at Peking University and Guolin Ke at DP Technology, targets a hard problem in pixel-space autoregressive image generation: directly predicting 16×16×3 pixel patches causes large single-step errors, while teacher-forced training creates train–inference mismatch and error accumulation. The method keeps a pixel-in, pixel-out autoregressive interface but internally predicts compact 16-dimensional intermediate states, then decodes them back to pixels with a learned causal pixel decoder. It also approximates inference-time rollout during training by constructing decoded, inference-like pixel inputs in parallel through the same state-to-pixel path used at test time, instead of expensive sequential on-policy sampling. On class-conditional ImageNet-1K at 256×256 resolution, PRA-S with 135M parameters reaches FID 2.58, beating the previous billion-scale pixel-space autoregressive result of 3.60, while PRA-L at 511M parameters improves to FID 1.94, a new state of the art for pixel-space AR models. Diagnostics show that using low-dimensional intermediate states and decoded pixel inputs is complementary: for 256×256 generation, direct high-dimensional AR drops to FID 7.68, input noise injection alone only reduces 9.94 to 7.68, and the PRA decoded-input path yields 2.88 in the ablation setting. Beyond generation, PRA-L also reaches 68.80% top-1 accuracy on ImageNet linear probing, outperforming both diffusion and latent-space AR baselines and indicating that end-to-end pixel-space autoregression can support stronger visual representations.

Original abstract

Pixel-space continuous-token autoregressive (AR) generation directly models images as sequences of raw pixel patches, avoiding discrete tokenization or a separately pretrained tokenizer. However, it faces coupled challenges: high-dimensional patch generation causes large single-step errors, and teacher-forced training creates a train--inference gap that makes these errors accumulate across AR steps. Existing fixes such as $x$-prediction and input noise injection only partially mitigate these issues. Exact rollout training better matches inference-time conditions, but is impractical due to prohibitively slow sequential sampling. We propose \emph{Parallel Rollout Approximation} (PRA), a scalable framework that addresses both challenges jointly. PRA generates low-dimensional intermediate states instead of high-dimensional pixel patches, then maps them back to pixel-space tokens with a pixel decoder, preserving a pixel-in, pixel-out AR interface. It also constructs inference-like pixel inputs through the same intermediate-state-to-pixel path used at inference, independently across positions, approximating the pixel-feedback interface encountered during inference-time rollout while retaining parallel teacher-forced training. On class-conditional ImageNet-1K generation at $256\times256$ resolution, PRA-S with 135M parameters achieves an FID of 2.58, surpassing the previous billion-scale pixel-space AR result of 3.60. Scaling to PRA-L with 511M parameters further improves FID to 1.94, establishing a new state of the art among pixel-space AR models. Beyond generation, PRA achieves higher ImageNet classification probing accuracy than other AR and diffusion baselines, suggesting its potential for unified pixel-space image generation and understanding.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis