Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation
AuthorsJiayi Xu, Di He, Guolin Ke
Resources
This paper makes pixel-by-pixel image generators train more like they infer, delivering a new state of the art without needing a separate tokenizer.
Key results
ImageNet-1K 256×256 class-conditional generation with 135M parameters
Prior billion-scale pixel-space autoregressive result on ImageNet-1K 256×256
ImageNet-1K 256×256 class-conditional generation with 511M parameters
Default compact internal state size used by PRA
ImageNet-1K probing accuracy for PRA-L
What the paper found
Parallel Rollout Approximation, developed by Jiayi Xu, Di He at Peking University and Guolin Ke at DP Technology, targets a hard problem in pixel-space autoregressive image generation: directly predicting 16×16×3 pixel patches causes large single-step errors, while teacher-forced training creates train–inference mismatch and error accumulation. The method keeps a pixel-in, pixel-out autoregressive interface but internally predicts compact 16-dimensional intermediate states, then decodes them back to pixels with a learned causal pixel decoder. It also approximates inference-time rollout during training by constructing decoded, inference-like pixel inputs in parallel through the same state-to-pixel path used at test time, instead of expensive sequential on-policy sampling. On class-conditional ImageNet-1K at 256×256 resolution, PRA-S with 135M parameters reaches FID 2.58, beating the previous billion-scale pixel-space autoregressive result of 3.60, while PRA-L at 511M parameters improves to FID 1.94, a new state of the art for pixel-space AR models. Diagnostics show that using low-dimensional intermediate states and decoded pixel inputs is complementary: for 256×256 generation, direct high-dimensional AR drops to FID 7.68, input noise injection alone only reduces 9.94 to 7.68, and the PRA decoded-input path yields 2.88 in the ablation setting. Beyond generation, PRA-L also reaches 68.80% top-1 accuracy on ImageNet linear probing, outperforming both diffusion and latent-space AR baselines and indicating that end-to-end pixel-space autoregression can support stronger visual representations.
Original abstract
Pixel-space continuous-token autoregressive (AR) generation directly models images as sequences of raw pixel patches, avoiding discrete tokenization or a separately pretrained tokenizer. However, it faces coupled challenges: high-dimensional patch generation causes large single-step errors, and teacher-forced training creates a train--inference gap that makes these errors accumulate across AR steps. Existing fixes such as $x$-prediction and input noise injection only partially mitigate these issues. Exact rollout training better matches inference-time conditions, but is impractical due to prohibitively slow sequential sampling. We propose \emph{Parallel Rollout Approximation} (PRA), a scalable framework that addresses both challenges jointly. PRA generates low-dimensional intermediate states instead of high-dimensional pixel patches, then maps them back to pixel-space tokens with a pixel decoder, preserving a pixel-in, pixel-out AR interface. It also constructs inference-like pixel inputs through the same intermediate-state-to-pixel path used at inference, independently across positions, approximating the pixel-feedback interface encountered during inference-time rollout while retaining parallel teacher-forced training. On class-conditional ImageNet-1K generation at $256\times256$ resolution, PRA-S with 135M parameters achieves an FID of 2.58, surpassing the previous billion-scale pixel-space AR result of 3.60. Scaling to PRA-L with 511M parameters further improves FID to 1.94, establishing a new state of the art among pixel-space AR models. Beyond generation, PRA achieves higher ImageNet classification probing accuracy than other AR and diffusion baselines, suggesting its potential for unified pixel-space image generation and understanding.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.