NTH

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

AuthorsWanshun Su, Yang Shi, Feihu Liu, Ziwen Yu, Yan Min, Zhuoran Zhang, Qixun Wang, Haotian Wang, Shixuan Liu, Yuanxing Zhang, Peng Wu, Chengfu Huo, Liang Ding

August 9, 2026 2 min read
Watch on YouTube
The one-line take

OmniPack makes omni-modal LLMs far more efficient by compressing redundant audio-visual tokens while preserving most of their task-relevant information.

Key results

98.0%
Performance at 25%/12.5% retention

Relative performance preserved by OmniPack on Qwen2.5-Omni-7B.

16.7%
FLOPs at 25%/12.5% retention

Fraction of original FLOPs used by OmniPack on Qwen2.5-Omni-7B.

95.6%
Performance at 15%/7.5% retention

Relative performance preserved under aggressive two-stage compression.

10.0%
FLOPs at 15%/7.5% retention

Fraction of original FLOPs used at the 15%/7.5% setting.

4.5
Prefill speedup

Fold speedup achieved at the 15%/7.5% retention setting.

92.9%
Performance at 10%/5% retention

Relative performance retained with only 6.8% of original FLOPs.

What the paper found

OmniPack, developed by researchers from Northwestern Polytechnical University, Peking University, Alibaba Group, and Tsinghua University, is a training-free framework for reducing the visual and audio token burden in omni-modal LLMs. Evaluated with Qwen2.5-Omni-7B, Qwen2.5-Omni-3B, and MiniCPM-o-2.6 on AVUT, WorldSense, DailyOmni, VideoMME, and LVOmniBench, it divides compression into two stages. Before the LLM, modality-specific importance scoring, global spatiotemporal coverage using DPC-KNN, and similarity-aware merging preserve salient and distributed evidence. After multimodal interaction inside the LLM, query-conditioned compression uses full-text guidance, audio-visual collaboration, and within-modality diversity to refine the remaining tokens. Against methods including OmniZip, OmniSIFT, and SEATS, OmniPack preserves 98.0% of original performance while using 16.7% of the original FLOPs on Qwen2.5-Omni-7B. At a 15%/7.5% pre-LLM/final retention setting, it retains 95.6% performance with 10.0% of original FLOPs and achieves a 4.5-fold prefill speedup; under 10%/5% retention, it still retains 92.9% performance at only 6.8% of original FLOPs. The results show that structural compression before semantic multimodal alignment and query-aware refinement afterward are complementary, enabling substantially more aggressive inference acceleration without retraining.

Original abstract

Omni-modal large language models (Omni-LLMs) have achieved remarkable performance on audio-visual understanding tasks, but processing long and highly redundant visual and audio token sequences incurs substantial computational overhead, demanding aggressive token compression for efficient deployment. Existing methods often degrade at low token budgets: pre-LLM compression may discard structurally important and globally distributed evidence, whereas inner-LLM compression often underexploits query-conditioned audio-visual collaboration. To address these limitations, we propose OmniPack, a training-free framework that coordinates structural compression before the LLM with task-relevant semantic refinement within the LLM. Before the LLM, OmniPack removes structural redundancy through modality-specific importance, global coverage, and similarity-aware merging. After sufficient multimodal interaction, it further consolidates diverse, task-relevant representations through textual guidance and audio-visual collaboration. Extensive experiments on five benchmarks with three Omni-LLM backbones demonstrate that OmniPack consistently achieves the best performance-efficiency trade-off across diverse retention ratios, outperforming all existing methods. Notably, on Qwen2.5-Omni-7B, OmniPack preserves 98.0% of the original performance while reducing FLOPs to 16.7%, and still retains 92.9% of the original performance with only 6.8% of the original FLOPs.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis