ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations
AuthorsJunke Wang, Xiao Wang, Jiacheng Pan, Xuefeng Hu, Feng Li, Jingxiang Sun, Chaorui Deng, Zilong Chen, Yunpeng Chen, Kaibin Tian, Matthew Gwilliam, Hao Chen, Danhui Guan, Kun Xu, Weilin Huang, Zuxuan Wu, Haoqi Fan, Yu-Gang Jiang, Zhenheng Yang
ARM is a multimodal AI model that turns images into discrete tokens so one autoregressive system can understand, generate, and edit visuals, with reinforcement learning improving both quality and instruction following.
Key results
Qwen2.5-based autoregressive multimodal model
Multimodal tokens used in the initial large-scale training stage
High-quality instruction-following tokens used for supervised fine-tuning
Multimodal understanding benchmark score
Image editing overall score after reinforcement learning
What the paper found
ARM, from Fudan University and ByteDance TikTok/ByteDance Seed, is an autoregressive large multimodal model that unifies image understanding, generation, and instruction-guided editing by predicting a single discrete token stream. Its key novelty is a unified visual tokenizer built on frozen SigLIP2-SO400M-512 features and Finite Scalar Quantization, trained with four complementary losses—captioning, pixel-space diffusion reconstruction, sigmoid contrastive alignment, and feature distillation—so the same token space preserves both semantics and visual detail. ARM then scales a 7B Qwen2.5 backbone on 2.5T multimodal tokens, adds supervised fine-tuning on 0.2B instruction tokens, and applies GRPO with GPT-o3 and GPT-4.1 as reward models to align generation and editing preferences. The model reports 40.2 on MMMU, 87.3 on POPE, 0.86 on GenEval, 0.56 on WISE after RL, and 6.68 on GEdit-Bench-EN G_O, while its RL stage produces clear cross-task synergy: WISE rises from 0.50 to 0.56 and GEdit-Bench-EN G_O from 5.75 to 6.68 without degrading understanding. The tokenizer ablation also shows why the design matters: adding semantic and feature regularization lifts ImageNet zero-shot accuracy from 0.2 to 80.2 and PSNR from 15.2 to 19.6, indicating that discrete representations can support both recognition and high-fidelity synthesis in one autoregressive framework.
Original abstract
This paper introduces ARM, a discrete representation-based AutoRegressive Model that unifies image understanding, generation, and editing within a next-token prediction framework. ARM is built on three efforts: first, we train a discrete semantic visual tokenizer that maps images into compact token sequences. Our tokenizer is supervised with multiple objectives that jointly promote semantic discriminability, language alignment and faithful reconstruction, thereby supporting diverse tasks in a shared latent space. With this, we train a 7B autoregressive model over large-scale text and image token sequences, seamlessly developing vision-language perception and generation capabilities. Finally, to further improve preference-aligned behavior for text-to-image generation and instruction-guided editing, ARM applies reinforcement learning (RL) to optimize task-level objectives such as visual quality, instruction adherence, and edit consistency. Surprisingly, the results show that RL not only substantially improves performance on the target tasks (e.g., raising WISE overall from 0.50 to 0.56, GEdit-Bench-EN G_O from 5.75 to 6.68), but also induces cross-task synergy between text-to-image generation and editing. Collectively, these findings highlight autoregressive modeling, when paired with strong representations and preference optimization, as a scalable foundation for multimodal intelligence. Code: https://github.com/wdrink/ARM.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.