Kirin: Animal Motion Generation from In-the-Wild Video
AuthorsBrian Nlong Zhao, Zhuoyang Pan, James M. Rehg, Jiajun Wu, Shangzhe Wu
Resources
Kirin turns large collections of animal videos into a multimodal system that can generate and animate realistic movements for diverse species.
Key results
Number of reconstructed in-the-wild animal motion sequences.
Six Gemini 2.5 Flash-generated descriptions per video sequence.
Held-out sequences used for benchmark evaluation.
AiM3D score for Kirin with text-plus-image conditioning.
AiM3D score for Kirin with text-plus-image conditioning.
Human-validated proportion of reconstructed motions judged physically plausible and smooth.
What the paper found
Kirin addresses the scarcity of realistic quadruped motion data by extracting temporally smooth 3D skeletal sequences from in-the-wild video. Starting with AiM footage across 23 quadruped categories, the authors refine SMAL-based pose estimates with keypoint alignment and rotation smoothness, recover global translation using SpatialTrackerV2, and pair each clip with six Gemini 2.5 Flash captions to create AiM3D: 29,979 motion sequences and 179,874 aligned descriptions, including a 230-sequence test split. Its generation model adapts MDM into a diffusion system conditioned jointly on text and an animal image, using frozen DistilBERT and DINOv3 encoders with classifier-free guidance, so the output respects both the requested action and the animal’s morphology. On AiM3D, text-plus-image conditioning reaches 0.043 top-1 R-Precision and 6.248 FID, outperforming AniMo variants; the model also generalizes to the external AnimalML3D benchmark. For production, Kirin uses Rodin for image-to-3D reconstruction, fits the SMAL skeleton, transfers skinning weights, and applies linear blend skinning to generate renderable animated meshes. Reconstruction runs in under 1 second per frame on an NVIDIA A40, while human review rates 86% of reconstructed motions as satisfactory. The main limitations are monocular depth ambiguity, occlusion-related leg errors, missing explicit tail dynamics, and a shared skeleton prior that may inadequately represent species with very different anatomies.
Original abstract
Understanding animal motion is fundamental to modeling animal behavior and biomechanics, yet progress in this area lags far behind human motion research due to the scarcity of high-quality motion data. While human motion can be captured in controlled environments, it is impractical for most animal species, resulting in small, domain-limited datasets that restrict downstream applications such as animation. To address this challenge, we introduce Kirin, a framework that reconstructs motion from video, learns motion priors at scale, and generates realistic motion that can be directly applied to animated assets. Using large collections of in-the-wild animal videos, we reconstruct 3D motion sequences and pair them with captions to create AiM3D, the first large-scale dataset offering aligned video-text-motion tuples for quadruped animals. Building on this dataset, we develop a visual-guided motion generation model that conditions on both text and image to guide the generation of realistic motion across diverse animal species. Finally, by leveraging an off-the-shelf image-to-3D model, we automatically rig and animate 3D meshes using generated motion, producing ready-to-render animated animals. Together, our dataset and framework establish a new foundation for large-scale, text and image conditioned animal motion generation and animation. Project page: https://kirin-ani.github.io/.
Read the original paperMore in Generative Models
Browse all 63 papers →RULER: Instance-aware Rubric Rewards for SVG Generation
Hangyu Ran, Yuhao Zheng, Yingying Zhang, Kevin Qinghong Lin, Han Peng
RULER uses instruction-specific visual rubrics as reinforcement-learning rewards to make SVG generation more faithful, stylish, and resistant to reward hacking.
Think Before You Score: Thinking Reward Model for Visual Generation
Xuehai Bai, Zhenchen Tang, Yang Shi, Dianyi Wang, Tengfei Liu, Wanshun Su, Xuanyu Zhu, Ruohui Wang, Haiwen Diao, Haotian Wang, Xiaoling Gu, Yuanxing Zhang
A visual reward model that first decides what matters in each image-generation case, then scores outputs with detailed rubrics to provide better training signals.
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation
Yubo Zhu, Yawen Shao, Ziyun Dai, Zixun Fang, Kai Zhu, Siyang Sun, Haolan Xue, Chuxin Wang, Tingyu Weng, Jingming Luo, Chen Shi, Lianghua Huang, Yufeng Ai, Yuzheng Wang, Wenyuan Zhang, Yu Shang, Yuxiang Bao, Zoubin Bi, Jie Xiao, Jinbo Xing, Jiaxing Zhao, Chongyang Zhong, Hengjian Chen, Chenwei Xie, Akide Liu, Zhehan Kan, Yu Liu, Wei Zhai, Sheng Zhong, Wei Tong
WanPE turns ordinary text prompts into director-level cinematic plans, substantially improving the quality and consistency of long-form AI-generated videos.