Channel-wise Vector Quantization
AuthorsWei Song, Tianhang Wang, Yitong Chen, Tong Zhang, Zuxuan Wu, Ming Li, Jiaqi Wang, Kaicheng Yu
Resources
This paper proposes a new way to turn images into discrete tokens by quantizing channels instead of patches, enabling a fresh autoregressive image generator that paints details progressively from coarse structure to fine texture.
Key results
CVQ with a 16,384-entry codebook
Conventional VQ at 16,384 entries
Vanilla CVQ reconstruction
Vanilla CVQ reconstruction
Largest codebook setting reported for CVQ with 96.1% utilization
What the paper found
Channel-wise Vector Quantization, from Shanghai Innovation Institute, Westlake University, Zhejiang University, Fudan University, JD.COM, and the University of Chinese Academy of Sciences, replaces patch-wise image tokens with channel-wise tokens: each latent channel becomes one discrete codeword, and autoregressive generation is reformulated as next-channel prediction rather than next-patch prediction. The key result is structural, not incremental: CVQ reaches 100% codebook utilization with a 16,384-entry codebook, while conventional VQ collapses to 4.5% utilization at the same size. On ImageNet-1K reconstruction, vanilla CVQ improves rFID to 2.60 at 256 tokens and 0.88 at 1024 tokens, with PSNR rising to 25.02 dB and SSIM to 0.723. The paper also shows that scaling the codebook to 65,536 entries preserves 96.1% utilization and reduces rFID to 2.32, a 52% reconstruction improvement over the VQ baseline. Built on CVQ, the channel-wise autoregressive model CAR is trained on 80M text-image pairs and a Qwen3-4B/8B backbone; the 8B version reaches GenEval 0.79 and DPG 86.72, while the 4B version scores 0.75 GenEval and 83.82 DPG. A simple nested channel-dropout scheme adds a coarse-to-fine ordering and boosts generation by 0.12 GenEval and 9.38 DPG without hurting reconstruction, indicating that channel tokens can support both efficient compression and strong text-to-image synthesis.
Original abstract
We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens. Unlike conventional vector quantization, which assigns a discrete token to each patch feature vector, CVQ quantizes each channel of the feature map. This formulation represents an image as discrete levels of visual details, rather than as a grid of spatial patches. Based on CVQ, we introduce a new visual autoregressive framework with "next-channel prediction". Instead of rendering images patch by patch in raster order, our Channel-wise Autoregressive (CAR) model predicts image channels sequentially, producing progressively enriched visual details. Specifically, it first sketches global structure and then refines fine-grained attributes, akin to a human artist's workflow. Empirically, we show that: (1) CVQ achieves 100% codebook utilization with a 16K+ codebook size without any bells and whistles, and substantially improves reconstruction quality over conventional VQ; and (2) CAR attains a DPG score of 86.7 and a GenEval score of 0.79, demonstrating strong effectiveness for text-to-image generation.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.