GEAR: Guided End-to-End AutoRegression for Image Synthesis
AuthorsBin Lin, Zheyuan Liu, Chenguo Lin, Sixiang Chen, Yunyang Ge, Yunlong Lin, Jianwei Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Li Yuan
Resources
GEAR lets an image generator and its tokenizer learn together, making autoregressive image synthesis train faster and produce better visual representations.
Key results
GEAR w/ CFG on ImageNet-1K at 256×256 for the 111M model
GEAR w/ CFG on ImageNet-1K at 256×256 for the 343M model
GEAR w/ CFG on ImageNet-1K at 256×256 for the 775M model
GEAR w/ CFG on GPIC after 390k steps
Image-text training set used for text-to-image experiments
Naive end-to-end training with straight-through estimator on ImageNet
What the paper found
GEAR, or Guided End-to-End AutoRegression, from Peking University and Tencent Hunyuan, reframes visual generation by training a vector-quantized tokenizer and an autoregressive generator together instead of freezing the tokenizer first. Its core technical move is a dual read-out of the discrete codebook assignment: a hard one-hot branch trains next-token prediction on the exact inference tokens, while a differentiable soft branch carries a representation-alignment loss back into the tokenizer, avoiding the collapse seen with straight-through estimation. On ImageNet-1K at 256×256, GEAR improves over LlamaGen-REPA at every scale: with classifier-free guidance, gFID drops from 6.00 to 4.95 for the 111M model, from 3.15 to 2.95 for 343M, and from 2.68 to 2.52 for 775M, while reaching up to 10× faster gFID convergence. The same tokenizer transfer works on text-to-image synthesis with GPIC’s 100M-image corpus and Qwen3-1.7B conditioning, where GEAR lowers FDD to 256.9, 177.4, 138.0, and 115.3 at 50k, 100k, 200k, and 390k steps, respectively. Ablations show the method is quantizer-agnostic across VQVAE, LFQ, and IBQ, and that replacing the soft bridge with STE catastrophically degrades generation to gFID 104.932 and rFID 59.723. Representation analysis finds an important reversal of diffusion-style alignment: the tokenizer becomes less DINOv2-like, but the AR hidden states become more patch-wise DINOv2-like and more spatially coherent, which explains the improved predictability and sample quality.
Original abstract
Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and then frozen, after which a generator is trained on its discrete indices or continuous latents. This decoupling leaves the tokenizer unaware of what the generator finds easy to model. We present GEAR (Guided End-to-end AutoRegression), which trains a vector-quantized (VQ) tokenizer and an autoregressive (AR) generator jointly and end-to-end, guided by representation alignment. The key obstacle is that the VQ index fed to the AR model is non-differentiable, so gradients cannot reach the tokenizer, and a straight-through estimator collapses. GEAR resolves this with a dual read-out of the codebook assignment. A hard, one-hot branch trains the AR with next-token prediction, while a differentiable soft branch carries a representation-alignment loss that flows back to guide only the tokenizer. The AR model thereby steers its tokenizer toward an index distribution it can predict more easily. This shifts the alignment burden from the tokenizer to the AR: the tokenizer's own features become less DINOv2-like while the AR's become more so, the opposite of diffusion-side recipes that make the latent itself semantic. GEAR speeds up ImageNet gFID convergence by up to 10x relative to the strong LlamaGen-REPA baseline, learns markedly better patch-level and spatially-coherent features, and generalizes across quantizers (VQVAE, LFQ, IBQ) and to text-to-image generation.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.