AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
AuthorsYiheng Li, Zhuo Li, Ruibing Hou, Yingjie Chen, Hong Chang, Hao Liu, Shiguang Shan
Resources
AnyMo introduces a large multimodal motion dataset and a masked-modeling transformer that can generate human motion from flexible combinations of inputs like text, speech, music, and trajectories.
Key results
high-quality motion sequences in the dataset
hours of motion data in the dataset
largest trained AnyMo variant
AnyMo-3B on OmniHuMo-Text test set
AnyMo-3B on OmniHuMo-Text test set
with trajectory conditioning in text-driven generation, measured in cm
What the paper found
AnyMo, from researchers at the Chinese Academy of Sciences and independent collaborators, tackles conditional human motion generation with a masked-modeling design that scales to arbitrary combinations of text, speech, music, and trajectory inputs. Its foundation is OmniHuMo, a new web-video-derived dataset with over 5,000 hours and 3.2 million motion sequences, each paired with text and a subset of about 500 hours also aligned with audio. The model replaces single-stream tokenization with a 4-layer Residual FSQ tokenizer using 2048 codes per layer, then trains a LLaMA-based bidirectional Transformer with parallel masked prediction across residual token streams. This lets AnyMo model global temporal context while conditioning on multiple modalities in a shared embedding space. On OmniHuMo-Text, scaling the model from 111M to 3B parameters steadily improves FID from 262.10 to 55.59 and raises R-Precision top-1 from 0.63 to 0.75. In multimodal conditioning, adding trajectory input cuts text-guided average trajectory error from 50.50 cm to 27.16 cm and lowers the share of errors above 50 cm from 0.52 to 0.33. The tokenizer also benefits from scale: reconstruction MPJPE drops from 94.55 mm at 0.05M data to 27.92 mm at 3M. Overall, the paper argues that large-scale multimodal alignment, not just larger models, is the key driver of flexible motion synthesis.
Original abstract
Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific architectures, leaving cross-modal interactions and the scaling laws of multimodal-conditioned synthesis largely underexplored. A key bottleneck is the scarcity of large-scale modality-aligned motion data, limiting generalization across diverse control signals. In this work, we introduce OmniHuMo, a large-scale, high-quality dataset comprising over 5,000 hours of motion and 3.2 million sequences with precisely aligned multimodal annotations (e.g., text, speech, music, and trajectory). Leveraging OmniHuMo, we propose AnyMo, a unified multimodal framework combining a Residual FSQ-based motion tokenizer with a scalable masked modeling transformer, enabling high-quality motion synthesis under arbitrary modality combinations. Extensive experiments show that AnyMo achieves high-fidelity synthesis while offering flexible control over both spatial and stylistic attributes.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.