NTH

ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations

AuthorsJunke Wang, Xiao Wang, Jiacheng Pan, Xuefeng Hu, Feng Li, Jingxiang Sun, Chaorui Deng, Zilong Chen, Yunpeng Chen, Kaibin Tian, Matthew Gwilliam, Hao Chen, Danhui Guan, Kun Xu, Weilin Huang, Zuxuan Wu, Haoqi Fan, Yu-Gang Jiang, Zhenheng Yang

June 22, 2026 2 min read
Watch on YouTube
The one-line take

ARM is a multimodal AI model that turns images into discrete tokens so one autoregressive system can understand, generate, and edit visuals, with reinforcement learning improving both quality and instruction following.

Key results

7B
Backbone size

Qwen2.5-based autoregressive multimodal model

2.5T
Pretraining tokens

Multimodal tokens used in the initial large-scale training stage

0.2B
SFT tokens

High-quality instruction-following tokens used for supervised fine-tuning

40.2
MMMU

Multimodal understanding benchmark score

6.68
GEdit-Bench-EN G_O

Image editing overall score after reinforcement learning

What the paper found

ARM, from Fudan University and ByteDance TikTok/ByteDance Seed, is an autoregressive large multimodal model that unifies image understanding, generation, and instruction-guided editing by predicting a single discrete token stream. Its key novelty is a unified visual tokenizer built on frozen SigLIP2-SO400M-512 features and Finite Scalar Quantization, trained with four complementary losses—captioning, pixel-space diffusion reconstruction, sigmoid contrastive alignment, and feature distillation—so the same token space preserves both semantics and visual detail. ARM then scales a 7B Qwen2.5 backbone on 2.5T multimodal tokens, adds supervised fine-tuning on 0.2B instruction tokens, and applies GRPO with GPT-o3 and GPT-4.1 as reward models to align generation and editing preferences. The model reports 40.2 on MMMU, 87.3 on POPE, 0.86 on GenEval, 0.56 on WISE after RL, and 6.68 on GEdit-Bench-EN G_O, while its RL stage produces clear cross-task synergy: WISE rises from 0.50 to 0.56 and GEdit-Bench-EN G_O from 5.75 to 6.68 without degrading understanding. The tokenizer ablation also shows why the design matters: adding semantic and feature regularization lifts ImageNet zero-shot accuracy from 0.2 to 80.2 and PSNR from 15.2 to 19.6, indicating that discrete representations can support both recognition and high-fidelity synthesis in one autoregressive framework.

Original abstract

This paper introduces ARM, a discrete representation-based AutoRegressive Model that unifies image understanding, generation, and editing within a next-token prediction framework. ARM is built on three efforts: first, we train a discrete semantic visual tokenizer that maps images into compact token sequences. Our tokenizer is supervised with multiple objectives that jointly promote semantic discriminability, language alignment and faithful reconstruction, thereby supporting diverse tasks in a shared latent space. With this, we train a 7B autoregressive model over large-scale text and image token sequences, seamlessly developing vision-language perception and generation capabilities. Finally, to further improve preference-aligned behavior for text-to-image generation and instruction-guided editing, ARM applies reinforcement learning (RL) to optimize task-level objectives such as visual quality, instruction adherence, and edit consistency. Surprisingly, the results show that RL not only substantially improves performance on the target tasks (e.g., raising WISE overall from 0.50 to 0.56, GEdit-Bench-EN G_O from 5.75 to 6.68), but also induces cross-task synergy between text-to-image generation and editing. Collectively, these findings highlight autoregressive modeling, when paired with strong representations and preference optimization, as a scalable foundation for multimodal intelligence. Code: https://github.com/wdrink/ARM.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis