RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
AuthorsSiyong Jian, Siyuan Li, Luyuan Zhang, Zedong Wang, Xin Jin, Ying Li, Cheng Tan, Huan Wang
Resources
RankE is a new way to fine-tune discrete text-to-image generators by updating both the token policy and decoder together, so better reward scores actually translate into better images.
Key results
AR backbone evaluated with RankE
RankE on LlamaGen-XL under CLIP reward
RankE on LlamaGen-XL under CLIP reward
RankE on Janus-Pro-1B under CLIP reward
RankE on Janus-Pro-1B under CLIP reward
Curated post-training corpus used for RankE
What the paper found
RankE, from Westlake University, Zhejiang University, Tsinghua University, Hong Kong University of Science and Technology, and Shanghai AI Lab, is a first end-to-end post-training framework for discrete text-to-image generation that targets a failure mode the paper calls Latent Covariate Shift: reward optimization improves the autoregressive policy while the frozen VQ decoder drifts out of its training distribution and image fidelity collapses. The method alternates between GRPO-based policy updates and decoder adaptation, using a ranking objective at two levels: token-level group-relative advantages for the policy and pixel-level reward-weighted Rank-GAN plus direct reward backpropagation for the decoder, while anchoring the decoder with reconstruction and EMA-consistency regularization. On LlamaGen-XL at 775M parameters, RankE breaks the usual CLIP–FID trade-off on MS-COCO 30K, reaching 33.76 CLIP and 15.21 FID, versus standard RL’s 32.45 CLIP and 17.76 FID; on Janus-Pro-1B it also improves CLIP to 33.86 and lowers FID to 25.19. The mechanism is supported by diagnostics showing standard RL widens token-distribution KL divergence by 24% and reduces codebook entropy, while RankE keeps both near the SFT regime. Ablations confirm that removing Rank-GAN, reconstruction, consistency, or reward backpropagation degrades performance, and the full method is trained on 8× NVIDIA A100 GPUs with a 16,384-entry codebook over a 15K-caption corpus.
Original abstract
Discrete autoregressive (AR) text-to-image (T2I) models pair a VQ tokenizer with an AR policy, and current post-training pipelines optimize only the policy while keeping the VQ decoder frozen. Recent diffusion T2I work, exemplified by REPA-E, has shown that the VAE itself constitutes a key alignment bottleneck, yet no analogous investigation exists for discrete AR models. We show that policy-only optimization induces Latent Covariate Shift: as the policy evolves, the resulting token distribution diverges from the ground-truth distribution on which the decoder was trained, such that reward scores improve while decoded image quality degrades. To address this mismatch, we propose RankE, the first end-to-end post-training framework for discrete T2I generation. Rather than optimizing the policy against a fixed decoder, RankE co-evolves both components through alternating optimization: each module maximizes a ranking-based alignment objective while being regularized by a stability-preserving anchor suited to its parameter space. This co-evolution breaks the fidelity--alignment trade-off that plagues frozen-decoder approaches: on LlamaGen-XL (775M), standard RL improves CLIP but degrades FID, whereas RankE improves both simultaneously (FID 15.21, CLIP 33.76 on MS-COCO 30K). Consistent gains on Janus-Pro (1B) confirm that decoder co-evolution reliably converts reward optimization into pixel-space quality improvements.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.