NTH

SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning

AuthorsWenhao Yan, Fengjia Guo, Zhuoyi Yang, Jie Tang

July 12, 2026 2 min read
Watch on YouTube
The one-line take

SCAIL-2 is a new end-to-end character animation system that learns motion transfer directly from video, using synthetic training data and preference tuning to produce more faithful animations.

Key results

59376
MotionPair-60K pairs

End-to-end motion-transfer dataset size reported in the appendix

3500
Pretraining steps

Full-finetuning steps for the Wan2.1-14B-I2V backbone

400
DPO post-training steps

Additional post-training steps after pretraining

5000
DPO temperature

Bias-Aware DPO temperature parameter

4.63
Video-Bench imaging quality

Full model score in the multi-character ablation on Video-Bench

What the paper found

SCAIL-2 from Tsinghua University and Z.ai reframes controlled character animation as a fully end-to-end in-context conditioning problem, bypassing brittle pose skeletons and masked-background intermediates that lose occlusion, interaction, and environment information. The paper unifies character image animation, character replacement, and multi-character interaction through a single DiT-based Wan2.1-14B-I2V backbone, using direct video-token concatenation, in-context mask conditioning, and mode-specific shifted RoPE to separate animation and replacement behavior without changing the visual context. To solve the paired-data bottleneck, it builds MotionPair-60K, a heterogeneous synthetic corpus of 59,376 end-to-end motion-transfer pairs with a 3:1 mix of animation and replacement data, generated through an agentic pipeline that uses SCAIL, Wan-Animate, MoCha, SAM3, Gemini, and Google DeepMind’s Nano Banana. The model is trained for 3,500 steps, then refined with 400 steps of Bias-Aware DPO, where preference pairs are constructed from fine-grained hand errors and optimized with a hand-region mask; this post-training uses LoRA rank 128 and a DPO temperature of 5000. Across Studio-Bench and X-Dance Benchmark, SCAIL-2 reports stronger cross-identity motion following, identity isolation, and environment integration than SCAIL, Wan-Animate, MoCha, and even proprietary Kling 3.0 in human preference tests, while the ablation on multi-character video-benchmark scores shows the full model reaching 4.63 imaging quality and 4.23 temporal consistency. The central contribution is not a new motion representation, but a scalable recipe for learning controlled animation directly from raw video context and synthetic preference correction.

Original abstract

Controlled character animation requires transferring motion from a driving sequence to a reference character. Prior works heavily rely on intermediate representations, including pose skeletons to represent motion or masked background to represent environment, which inevitably leads to information loss. To address this, we present SCAIL-2, an framework that bypasses those intermediates and achieves \textbf{end-to-end} character animation. By directly concatenating driving videos to the sequence, the model can obtain all the required visual information from the input video. To address lack of end-to-end data, we unify sub-tasks of character animation with decoupled conditions and then curate a pipeline to synthesize MotionPair-60K, an end-to-end motion transfer dataset containing heterogeneous tasks of character animation. To archive the unification, we utilize in-context mask conditioning and mode-specific RoPE as soft guidance beyond textual instructions and raw visual information. To address synthetic discrepancy in detailed regions, we propose Bias-Aware DPO to construct preference items to mitigate the errors. Extensive experiments demonstrate that our method substantially outperforms existing state-of-the-art approaches in various character animation tasks. A large subset of synthetic data as well as model weights will be released at our project page: https://teal024.github.io/SCAIL-2/.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis