NTH

VibeToken: Scaling 1D Image Tokenizers and Autoregressive Models for Dynamic Resolution Generations

AuthorsMaitreya Patel, Jingtao Li, Weiming Zhuang, Yezhou Yang, Lingjuan Lv

June 9, 2026 2 min read
Watch on YouTube
The one-line take

VibeToken is a new image tokenization and generation approach that lets autoregressive models create high-resolution images at arbitrary aspect ratios with far less compute.

Key results

32–256
Tokenizer tokens

VibeToken encodes images with a dynamic token budget

1k
ImageNet-1k

Tokenizer and generator are trained on ImageNet-1k

0.40
rFID at 256×256

Large VibeToken reconstruction quality on ImageNet

2.40
rFID at 1024×1024

Large VibeToken reconstruction quality at high resolution

179G
Inference FLOPs

VibeToken-Gen keeps constant inference compute across resolutions

3.54
1024×1024 gFID

Best VibeToken-Gen generation result at high resolution

What the paper found

VibeToken, from SonyAI and Arizona State University, reframes image tokenization as a resolution-agnostic 1D Transformer problem so autoregressive generation can scale to arbitrary aspect ratios without exploding token length. The tokenizer encodes images into a controllable 32–256 token sequence using dynamic grid position embeddings, adaptive patch embedding, adaptive decoder resolution, and length-uniform training on ImageNet-1k; in reconstruction it reaches 0.40 rFID at 256×256, 0.51 at 512×512, and 2.40 at 1024×1024 with the large variant, while preserving native 4× super-resolution without a separate upsampler. Building on these tokens, VibeToken-Gen is a class-conditioned AR generator that trains on ImageNet-1k and keeps inference compute fixed at 179 GFLOPs for any resolution, compared with about 11 TFLOPs for LlamaGen at 1024×1024. The strongest generator uses only 64 tokens, producing 1024×1024 images at 3.54 gFID in 0.46 s, versus NiT’s 5.87 gFID in 1.08 s, and it improves generation efficiency by 63.4× over the fixed-resolution AR baseline. The paper’s key novelty is that compute is decoupled from pixel count: VibeToken shifts the scaling burden from the generator to the tokenizer, making AR visual synthesis resolution-generalist rather than resolution-bound.

Original abstract

We introduce an efficient, resolution-agnostic autoregressive (AR) image synthesis approach that generalizes to arbitrary resolutions and aspect ratios, narrowing the gap to diffusion models at scale. At its core is VibeToken, a novel resolution-agnostic 1D Transformer-based image tokenizer that encodes images into a dynamic, user-controllable sequence of 32-256 tokens, achieving a state-of-the-art efficiency and performance trade-off. Building on VibeToken, we present VibeToken-Gen, a class-conditioned AR generator with out-of-the-box support for arbitrary resolutions while requiring significantly fewer compute resources. Notably, VibeToken-Gen synthesizes 1024x1024 images using only 64 tokens and achieves 3.94 gFID; by comparison, a diffusion-based state-of-the-art alternative requires 1,024 tokens and attains 5.87 gFID. In contrast to fixed-resolution AR models such as LlamaGen -- whose inference FLOPs grow quadratically with resolution (11T FLOPs at 1024x1024) -- VibeToken-Gen maintains a constant 179G FLOPs (63.4x efficient) independent of resolution. We hope VibeToken can help unlock the wide adoption of AR visual generative models in production use cases.

Read the original paper

More in Generative Models

Browse all 63 papers →
01Generative Model

RULER: Instance-aware Rubric Rewards for SVG Generation

Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng

RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.

Read analysis
02Generative Model

Think Before You Score: Thinking Reward Model for Visual Generation

Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang

A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.

Read analysis
03Generative Model

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong

WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.

Read analysis