NTH

Channel-wise Vector Quantization

AuthorsWei Song, Tianhang Wang, Yitong Chen, Tong Zhang, Zuxuan Wu, Ming Li, Jiaqi Wang, Kaicheng Yu

June 12, 2026 2 min read
Watch on YouTube
The one-line take

This paper proposes a new way to turn images into discrete tokens by quantizing channels instead of patches, enabling a fresh autoregressive image generator that paints details progressively from coarse structure to fine texture.

Key results

100%
codebook utilization

CVQ with a 16,384-entry codebook

4.5%
codebook utilization baseline

Conventional VQ at 16,384 entries

2.60
ImageNet-1K rFID (256 tokens)

Vanilla CVQ reconstruction

0.88
ImageNet-1K rFID (1024 tokens)

Vanilla CVQ reconstruction

65536
codebook size scale

Largest codebook setting reported for CVQ with 96.1% utilization

What the paper found

Channel-wise Vector Quantization, from Shanghai Innovation Institute, Westlake University, Zhejiang University, Fudan University, JD.COM, and the University of Chinese Academy of Sciences, replaces patch-wise image tokens with channel-wise tokens: each latent channel becomes one discrete codeword, and autoregressive generation is reformulated as next-channel prediction rather than next-patch prediction. The key result is structural, not incremental: CVQ reaches 100% codebook utilization with a 16,384-entry codebook, while conventional VQ collapses to 4.5% utilization at the same size. On ImageNet-1K reconstruction, vanilla CVQ improves rFID to 2.60 at 256 tokens and 0.88 at 1024 tokens, with PSNR rising to 25.02 dB and SSIM to 0.723. The paper also shows that scaling the codebook to 65,536 entries preserves 96.1% utilization and reduces rFID to 2.32, a 52% reconstruction improvement over the VQ baseline. Built on CVQ, the channel-wise autoregressive model CAR is trained on 80M text-image pairs and a Qwen3-4B/8B backbone; the 8B version reaches GenEval 0.79 and DPG 86.72, while the 4B version scores 0.75 GenEval and 83.82 DPG. A simple nested channel-dropout scheme adds a coarse-to-fine ordering and boosts generation by 0.12 GenEval and 9.38 DPG without hurting reconstruction, indicating that channel tokens can support both efficient compression and strong text-to-image synthesis.

Original abstract

We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens. Unlike conventional vector quantization, which assigns a discrete token to each patch feature vector, CVQ quantizes each channel of the feature map. This formulation represents an image as discrete levels of visual details, rather than as a grid of spatial patches. Based on CVQ, we introduce a new visual autoregressive framework with "next-channel prediction". Instead of rendering images patch by patch in raster order, our Channel-wise Autoregressive (CAR) model predicts image channels sequentially, producing progressively enriched visual details. Specifically, it first sketches global structure and then refines fine-grained attributes, akin to a human artist's workflow. Empirically, we show that: (1) CVQ achieves 100% codebook utilization with a 16K+ codebook size without any bells and whistles, and substantially improves reconstruction quality over conventional VQ; and (2) CAR attains a DPG score of 86.7 and a GenEval score of 0.79, demonstrating strong effectiveness for text-to-image generation.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis