NTH

RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution

AuthorsSiyong Jian, Siyuan Li, Luyuan Zhang, Zedong Wang, Xin Jin, Ying Li, Cheng Tan, Huan Wang

June 28, 2026 2 min read
Watch on YouTube
The one-line take

RankE is a new way to fine-tune discrete text-to-image generators by updating both the token policy and decoder together, so better reward scores actually translate into better images.

Key results

775M
LlamaGen-XL size

AR backbone evaluated with RankE

15.21
MS-COCO 30K FID

RankE on LlamaGen-XL under CLIP reward

33.76
MS-COCO 30K CLIP

RankE on LlamaGen-XL under CLIP reward

33.86
Janus-Pro CLIP

RankE on Janus-Pro-1B under CLIP reward

25.19
Janus-Pro FID

RankE on Janus-Pro-1B under CLIP reward

15K
Training set size

Curated post-training corpus used for RankE

What the paper found

RankE, from Westlake University, Zhejiang University, Tsinghua University, Hong Kong University of Science and Technology, and Shanghai AI Lab, is a first end-to-end post-training framework for discrete text-to-image generation that targets a failure mode the paper calls Latent Covariate Shift: reward optimization improves the autoregressive policy while the frozen VQ decoder drifts out of its training distribution and image fidelity collapses. The method alternates between GRPO-based policy updates and decoder adaptation, using a ranking objective at two levels: token-level group-relative advantages for the policy and pixel-level reward-weighted Rank-GAN plus direct reward backpropagation for the decoder, while anchoring the decoder with reconstruction and EMA-consistency regularization. On LlamaGen-XL at 775M parameters, RankE breaks the usual CLIP–FID trade-off on MS-COCO 30K, reaching 33.76 CLIP and 15.21 FID, versus standard RL’s 32.45 CLIP and 17.76 FID; on Janus-Pro-1B it also improves CLIP to 33.86 and lowers FID to 25.19. The mechanism is supported by diagnostics showing standard RL widens token-distribution KL divergence by 24% and reduces codebook entropy, while RankE keeps both near the SFT regime. Ablations confirm that removing Rank-GAN, reconstruction, consistency, or reward backpropagation degrades performance, and the full method is trained on 8× NVIDIA A100 GPUs with a 16,384-entry codebook over a 15K-caption corpus.

Original abstract

Discrete autoregressive (AR) text-to-image (T2I) models pair a VQ tokenizer with an AR policy, and current post-training pipelines optimize only the policy while keeping the VQ decoder frozen. Recent diffusion T2I work, exemplified by REPA-E, has shown that the VAE itself constitutes a key alignment bottleneck, yet no analogous investigation exists for discrete AR models. We show that policy-only optimization induces Latent Covariate Shift: as the policy evolves, the resulting token distribution diverges from the ground-truth distribution on which the decoder was trained, such that reward scores improve while decoded image quality degrades. To address this mismatch, we propose RankE, the first end-to-end post-training framework for discrete T2I generation. Rather than optimizing the policy against a fixed decoder, RankE co-evolves both components through alternating optimization: each module maximizes a ranking-based alignment objective while being regularized by a stability-preserving anchor suited to its parameter space. This co-evolution breaks the fidelity--alignment trade-off that plagues frozen-decoder approaches: on LlamaGen-XL (775M), standard RL improves CLIP but degrades FID, whereas RankE improves both simultaneously (FID 15.21, CLIP 33.76 on MS-COCO 30K). Consistent gains on Janus-Pro (1B) confirm that decoder co-evolution reliably converts reward optimization into pixel-space quality improvements.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis