NTH

GEAR: Guided End-to-End AutoRegression for Image Synthesis

AuthorsBin Lin, Zheyuan Liu, Chenguo Lin, Sixiang Chen, Yunyang Ge, Yunlong Lin, Jianwei Zhang, Miles Yang, Zhao Zhong, Liefeng Bo, Li Yuan

July 11, 2026 2 min read
Watch on YouTube
The one-line take

GEAR lets an image generator and its tokenizer learn together, making autoregressive image synthesis train faster and produce better visual representations.

Key results

4.95
ImageNet gFID improvement

GEAR w/ CFG on ImageNet-1K at 256×256 for the 111M model

2.95
ImageNet gFID improvement

GEAR w/ CFG on ImageNet-1K at 256×256 for the 343M model

2.52
ImageNet gFID improvement

GEAR w/ CFG on ImageNet-1K at 256×256 for the 775M model

115.3
GPIC FDD

GEAR w/ CFG on GPIC after 390k steps

100M
GPIC corpus size

Image-text training set used for text-to-image experiments

104.932
STE failure gFID

Naive end-to-end training with straight-through estimator on ImageNet

What the paper found

GEAR, or Guided End-to-End AutoRegression, from Peking University and Tencent Hunyuan, reframes visual generation by training a vector-quantized tokenizer and an autoregressive generator together instead of freezing the tokenizer first. Its core technical move is a dual read-out of the discrete codebook assignment: a hard one-hot branch trains next-token prediction on the exact inference tokens, while a differentiable soft branch carries a representation-alignment loss back into the tokenizer, avoiding the collapse seen with straight-through estimation. On ImageNet-1K at 256×256, GEAR improves over LlamaGen-REPA at every scale: with classifier-free guidance, gFID drops from 6.00 to 4.95 for the 111M model, from 3.15 to 2.95 for 343M, and from 2.68 to 2.52 for 775M, while reaching up to 10× faster gFID convergence. The same tokenizer transfer works on text-to-image synthesis with GPIC’s 100M-image corpus and Qwen3-1.7B conditioning, where GEAR lowers FDD to 256.9, 177.4, 138.0, and 115.3 at 50k, 100k, 200k, and 390k steps, respectively. Ablations show the method is quantizer-agnostic across VQVAE, LFQ, and IBQ, and that replacing the soft bridge with STE catastrophically degrades generation to gFID 104.932 and rFID 59.723. Representation analysis finds an important reversal of diffusion-style alignment: the tokenizer becomes less DINOv2-like, but the AR hidden states become more patch-wise DINOv2-like and more spatially coherent, which explains the improved predictability and sample quality.

Original abstract

Visual generative models are typically trained in two stages. A tokenizer is first trained for reconstruction and then frozen, after which a generator is trained on its discrete indices or continuous latents. This decoupling leaves the tokenizer unaware of what the generator finds easy to model. We present GEAR (Guided End-to-end AutoRegression), which trains a vector-quantized (VQ) tokenizer and an autoregressive (AR) generator jointly and end-to-end, guided by representation alignment. The key obstacle is that the VQ index fed to the AR model is non-differentiable, so gradients cannot reach the tokenizer, and a straight-through estimator collapses. GEAR resolves this with a dual read-out of the codebook assignment. A hard, one-hot branch trains the AR with next-token prediction, while a differentiable soft branch carries a representation-alignment loss that flows back to guide only the tokenizer. The AR model thereby steers its tokenizer toward an index distribution it can predict more easily. This shifts the alignment burden from the tokenizer to the AR: the tokenizer's own features become less DINOv2-like while the AR's become more so, the opposite of diffusion-side recipes that make the latent itself semantic. GEAR speeds up ImageNet gFID convergence by up to 10x relative to the strong LlamaGen-REPA baseline, learns markedly better patch-level and spatially-coherent features, and generalizes across quantizers (VQVAE, LFQ, IBQ) and to text-to-image generation.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis