SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning
AuthorsWenhao Yan, Fengjia Guo, Zhuoyi Yang, Jie Tang
Resources
SCAIL-2 is a new end-to-end character animation system that learns motion transfer directly from video, using synthetic training data and preference tuning to produce more faithful animations.
Key results
End-to-end motion-transfer dataset size reported in the appendix
Full-finetuning steps for the Wan2.1-14B-I2V backbone
Additional post-training steps after pretraining
Bias-Aware DPO temperature parameter
Full model score in the multi-character ablation on Video-Bench
What the paper found
SCAIL-2 from Tsinghua University and Z.ai reframes controlled character animation as a fully end-to-end in-context conditioning problem, bypassing brittle pose skeletons and masked-background intermediates that lose occlusion, interaction, and environment information. The paper unifies character image animation, character replacement, and multi-character interaction through a single DiT-based Wan2.1-14B-I2V backbone, using direct video-token concatenation, in-context mask conditioning, and mode-specific shifted RoPE to separate animation and replacement behavior without changing the visual context. To solve the paired-data bottleneck, it builds MotionPair-60K, a heterogeneous synthetic corpus of 59,376 end-to-end motion-transfer pairs with a 3:1 mix of animation and replacement data, generated through an agentic pipeline that uses SCAIL, Wan-Animate, MoCha, SAM3, Gemini, and Google DeepMind’s Nano Banana. The model is trained for 3,500 steps, then refined with 400 steps of Bias-Aware DPO, where preference pairs are constructed from fine-grained hand errors and optimized with a hand-region mask; this post-training uses LoRA rank 128 and a DPO temperature of 5000. Across Studio-Bench and X-Dance Benchmark, SCAIL-2 reports stronger cross-identity motion following, identity isolation, and environment integration than SCAIL, Wan-Animate, MoCha, and even proprietary Kling 3.0 in human preference tests, while the ablation on multi-character video-benchmark scores shows the full model reaching 4.63 imaging quality and 4.23 temporal consistency. The central contribution is not a new motion representation, but a scalable recipe for learning controlled animation directly from raw video context and synthetic preference correction.
Original abstract
Controlled character animation requires transferring motion from a driving sequence to a reference character. Prior works heavily rely on intermediate representations, including pose skeletons to represent motion or masked background to represent environment, which inevitably leads to information loss. To address this, we present SCAIL-2, an framework that bypasses those intermediates and achieves \textbf{end-to-end} character animation. By directly concatenating driving videos to the sequence, the model can obtain all the required visual information from the input video. To address lack of end-to-end data, we unify sub-tasks of character animation with decoupled conditions and then curate a pipeline to synthesize MotionPair-60K, an end-to-end motion transfer dataset containing heterogeneous tasks of character animation. To archive the unification, we utilize in-context mask conditioning and mode-specific RoPE as soft guidance beyond textual instructions and raw visual information. To address synthetic discrepancy in detailed regions, we propose Bias-Aware DPO to construct preference items to mitigate the errors. Extensive experiments demonstrate that our method substantially outperforms existing state-of-the-art approaches in various character animation tasks. A large subset of synthetic data as well as model weights will be released at our project page: https://teal024.github.io/SCAIL-2/.
Read the original paperMore in Computer Vision
Browse all 58 papers →All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts
Xingsong Ye, Yongkun Du, Jiaxin Zhang, Zhixian Li, Chong Sun, Chen Li, Jing Lyu, Lianwen Jin, Zhineng Chen
A lightweight script-aware mixture-of-experts model brings more accurate, scalable multilingual scene text recognition to many languages and scripts.
DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany
DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.
OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?
Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang
This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.