HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers
AuthorsGuozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang
Resources
HYDRA-X is a new unified multimodal model that uses one visual tokenizer for both images and videos, improving understanding, generation, and editing by doing more of the work inside the tokenizer itself.
Key results
HYDRA-X is instantiated at 7B scale
HYDRA-XTok Stage 3 reconstruction on ImageNet
HYDRA-XTok Stage 3 video reconstruction on DAVIS
HYDRA-XTok Stage 3 video reconstruction on UCF
HYDRA-X improves over Show-o2 on VBench Total
Tokenizer-stage source-target interaction improves source reconstruction consistency
What the paper found
HYDRA-X from Nanjing University, CASIA, Tencent Hunyuan, and Shanghai AI Lab is presented as the first native unified multimodal model to unify image and video tokenization inside a single Vision Transformer, building on the earlier HYDRA framework and Qwen2.5-7B-Instruct. Its tokenizer, HYDRA-XTok, uses a counterintuitive but effective recipe: frame-level causal tubelet attention instead of full spatiotemporal attention, hierarchical 2×2 temporal patchify instead of one-step 4× compression, and a lightweight Decompressor that restores compressed video latents for dual supervision from SigLIP-SO400M-patch16-naflex and InternVideo-Next-L. On reconstruction, the tokenizer reaches 32.04 PSNR on ImageNet, 28.19 PSNR on DAVIS, and 36.88 PSNR on UCF, while the fully trained 7B model improves over Show-o2 on video generation VBench by +1.87 QS, +3.26 SS, and +2.15 Total. For image understanding, it posts 86.5 on AI2D and 2350.0 on MME; for video understanding, it scores 59.1 on MVBench and 60.0 on Video-MME. The paper also shows that moving source-target interaction for image editing into the tokenizer, rather than the LLM, raises ImgEdit-Bench overall from 2.80 to 3.20 and source-reconstruction PSNR from 20.74 to 27.65, confirming that latent-level coupling is crucial for identity-faithful edits.
Original abstract
Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDRA-X, the first UMM that unifies image and video tokenization within a single Vision Transformer (ViT). Our design is driven by two core challenges: efficiently injecting spatiotemporal reconstruction capability into a native ViT, and embedding image- and video-level semantic awareness into the latent space. To address the first, comprehensive ablations reveal two key findings: (1) frame-level causal temporal attention suffices for visual reconstruction, whereas full spatiotemporal attention degrades it; and (2) hierarchical temporal compression substantially outperforms single-step alternatives. To tackle the second, we propose a lightweight decompressor that upsamples temporally compressed features under joint image-video teacher supervision, thereby enforcing complementary semantic structures within the compact latent space. Building on this holistic tokenizer, we further propose a principled improvement of the editing pipeline: source-target interaction should occur at the latent level inside the tokenizer rather than at the semantic level inside the LLM, substantially improving editing consistency and accelerating convergence. Instantiated at the 7B dense model, HYDRA-X achieves strong performance across image and video understanding and generation tasks, paving the way for future unified-tokenizer UMMs.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.