NTH

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

AuthorsGuozhen Zhang, Xuerui Qiu, Yutao Cui, Tianhui Song, Changlin Li, Junzhe Li, Tao Huang, Xiao Zhang, Yang Li, Jianbing Wu, Miles Yang, Zhao Zhong, Liefeng Bo, Limin Wang

June 22, 2026 2 min read
Watch on YouTube
The one-line take

HYDRA-X is a new unified multimodal model that uses one visual tokenizer for both images and videos, improving understanding, generation, and editing by doing more of the work inside the tokenizer itself.

Key results

7B
model scale

HYDRA-X is instantiated at 7B scale

32.04
ImageNet PSNR

HYDRA-XTok Stage 3 reconstruction on ImageNet

28.19
DAVIS PSNR

HYDRA-XTok Stage 3 video reconstruction on DAVIS

36.88
UCF PSNR

HYDRA-XTok Stage 3 video reconstruction on UCF

2.15
VBench Total gain

HYDRA-X improves over Show-o2 on VBench Total

27.65
ImgEdit source PSNR

Tokenizer-stage source-target interaction improves source reconstruction consistency

What the paper found

HYDRA-X from Nanjing University, CASIA, Tencent Hunyuan, and Shanghai AI Lab is presented as the first native unified multimodal model to unify image and video tokenization inside a single Vision Transformer, building on the earlier HYDRA framework and Qwen2.5-7B-Instruct. Its tokenizer, HYDRA-XTok, uses a counterintuitive but effective recipe: frame-level causal tubelet attention instead of full spatiotemporal attention, hierarchical 2×2 temporal patchify instead of one-step 4× compression, and a lightweight Decompressor that restores compressed video latents for dual supervision from SigLIP-SO400M-patch16-naflex and InternVideo-Next-L. On reconstruction, the tokenizer reaches 32.04 PSNR on ImageNet, 28.19 PSNR on DAVIS, and 36.88 PSNR on UCF, while the fully trained 7B model improves over Show-o2 on video generation VBench by +1.87 QS, +3.26 SS, and +2.15 Total. For image understanding, it posts 86.5 on AI2D and 2350.0 on MME; for video understanding, it scores 59.1 on MVBench and 60.0 on Video-MME. The paper also shows that moving source-target interaction for image editing into the tokenizer, rather than the LLM, raises ImgEdit-Bench overall from 2.80 to 3.20 and source-reconstruction PSNR from 20.74 to 27.65, confirming that latent-level coupling is crucial for identity-faithful edits.

Original abstract

Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDRA-X, the first UMM that unifies image and video tokenization within a single Vision Transformer (ViT). Our design is driven by two core challenges: efficiently injecting spatiotemporal reconstruction capability into a native ViT, and embedding image- and video-level semantic awareness into the latent space. To address the first, comprehensive ablations reveal two key findings: (1) frame-level causal temporal attention suffices for visual reconstruction, whereas full spatiotemporal attention degrades it; and (2) hierarchical temporal compression substantially outperforms single-step alternatives. To tackle the second, we propose a lightweight decompressor that upsamples temporally compressed features under joint image-video teacher supervision, thereby enforcing complementary semantic structures within the compact latent space. Building on this holistic tokenizer, we further propose a principled improvement of the editing pipeline: source-target interaction should occur at the latent level inside the tokenizer rather than at the semantic level inside the LLM, substantially improving editing consistency and accelerating convergence. Instantiated at the 7B dense model, HYDRA-X achieves strong performance across image and video understanding and generation tasks, paving the way for future unified-tokenizer UMMs.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis