NTH

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling

AuthorsYiheng Li, Zhuo Li, Ruibing Hou, Yingjie Chen, Hong Chang, Hao Liu, Shiguang Shan

June 18, 2026 2 min read
Watch on YouTube
The one-line take

AnyMo introduces a large multimodal motion dataset and a masked-modeling transformer that can generate human motion from flexible combinations of inputs like text, speech, music, and trajectories.

Key results

3.2M
OmniHuMo sequences

high-quality motion sequences in the dataset

5000
OmniHuMo duration

hours of motion data in the dataset

3B
AnyMo max model size

largest trained AnyMo variant

55.59
Text-driven FID

AnyMo-3B on OmniHuMo-Text test set

0.75
Text-driven R@1

AnyMo-3B on OmniHuMo-Text test set

27.16
Trajectory average error

with trajectory conditioning in text-driven generation, measured in cm

What the paper found

AnyMo, from researchers at the Chinese Academy of Sciences and independent collaborators, tackles conditional human motion generation with a masked-modeling design that scales to arbitrary combinations of text, speech, music, and trajectory inputs. Its foundation is OmniHuMo, a new web-video-derived dataset with over 5,000 hours and 3.2 million motion sequences, each paired with text and a subset of about 500 hours also aligned with audio. The model replaces single-stream tokenization with a 4-layer Residual FSQ tokenizer using 2048 codes per layer, then trains a LLaMA-based bidirectional Transformer with parallel masked prediction across residual token streams. This lets AnyMo model global temporal context while conditioning on multiple modalities in a shared embedding space. On OmniHuMo-Text, scaling the model from 111M to 3B parameters steadily improves FID from 262.10 to 55.59 and raises R-Precision top-1 from 0.63 to 0.75. In multimodal conditioning, adding trajectory input cuts text-guided average trajectory error from 50.50 cm to 27.16 cm and lowers the share of errors above 50 cm from 0.52 to 0.33. The tokenizer also benefits from scale: reconstruction MPJPE drops from 94.55 mm at 0.05M data to 27.92 mm at 3M. Overall, the paper argues that large-scale multimodal alignment, not just larger models, is the key driver of flexible motion synthesis.

Original abstract

Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific architectures, leaving cross-modal interactions and the scaling laws of multimodal-conditioned synthesis largely underexplored. A key bottleneck is the scarcity of large-scale modality-aligned motion data, limiting generalization across diverse control signals. In this work, we introduce OmniHuMo, a large-scale, high-quality dataset comprising over 5,000 hours of motion and 3.2 million sequences with precisely aligned multimodal annotations (e.g., text, speech, music, and trajectory). Leveraging OmniHuMo, we propose AnyMo, a unified multimodal framework combining a Residual FSQ-based motion tokenizer with a scalable masked modeling transformer, enabling high-quality motion synthesis under arbitrary modality combinations. Extensive experiments show that AnyMo achieves high-fidelity synthesis while offering flexible control over both spatial and stylistic attributes.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis