VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations
AuthorsMaitreya Patel, Jingtao Li, Weiming Zhuang, Yezhou Yang, Lingjuan Lv
Resources
VibeToken is a new image tokenization and generation approach that lets autoregressive models create high-resolution images at arbitrary aspect ratios with far less compute.
Key results
VibeToken encodes images with a dynamic token budget
Tokenizer and generator are trained on ImageNet-1k
Large VibeToken reconstruction quality on ImageNet
Large VibeToken reconstruction quality at high resolution
VibeToken-Gen keeps constant inference compute across resolutions
Best VibeToken-Gen generation result at high resolution
What the paper found
VibeToken, from SonyAI and Arizona State University, reframes image tokenization as a resolution-agnostic 1D Transformer problem so autoregressive generation can scale to arbitrary aspect ratios without exploding token length. The tokenizer encodes images into a controllable 32–256 token sequence using dynamic grid position embeddings, adaptive patch embedding, adaptive decoder resolution, and length-uniform training on ImageNet-1k; in reconstruction it reaches 0.40 rFID at 256×256, 0.51 at 512×512, and 2.40 at 1024×1024 with the large variant, while preserving native 4× super-resolution without a separate upsampler. Building on these tokens, VibeToken-Gen is a class-conditioned AR generator that trains on ImageNet-1k and keeps inference compute fixed at 179 GFLOPs for any resolution, compared with about 11 TFLOPs for LlamaGen at 1024×1024. The strongest generator uses only 64 tokens, producing 1024×1024 images at 3.54 gFID in 0.46 s, versus NiT’s 5.87 gFID in 1.08 s, and it improves generation efficiency by 63.4× over the fixed-resolution AR baseline. The paper’s key novelty is that compute is decoupled from pixel count: VibeToken shifts the scaling burden from the generator to the tokenizer, making AR visual synthesis resolution-generalist rather than resolution-bound.
Original abstract
We introduce an efficient, resolution-agnostic autoregressive (AR) image synthesis approach that generalizes to arbitrary resolutions and aspect ratios, narrowing the gap to diffusion models at scale. At its core is VibeToken, a novel resolution-agnostic 1D Transformer-based image tokenizer that encodes images into a dynamic, user-controllable sequence of 32-256 tokens, achieving a state-of-the-art efficiency and performance trade-off. Building on VibeToken, we present VibeToken-Gen, a class-conditioned AR generator with out-of-the-box support for arbitrary resolutions while requiring significantly fewer compute resources. Notably, VibeToken-Gen synthesizes 1024x1024 images using only 64 tokens and achieves 3.94 gFID; by comparison, a diffusion-based state-of-the-art alternative requires 1,024 tokens and attains 5.87 gFID. In contrast to fixed-resolution AR models such as LlamaGen -- whose inference FLOPs grow quadratically with resolution (11T FLOPs at 1024x1024) -- VibeToken-Gen maintains a constant 179G FLOPs (63.4x efficient) independent of resolution. We hope VibeToken can help unlock the wide adoption of AR visual generative models in production use cases.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.