NTH

AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model

AuthorsKwan Yun, Serin Yoon, Sunjin Jung, Jung Eun Yoo, Inyup Lee, Junyong Noh

August 26, 2026 2 min read
Watch on YouTube
The one-line take

AnyTalk turns generic video-generation knowledge into real-time, lip-synced 3D animation for almost any character without requiring character-specific motion data.

Key results

5
Character coverage

Number of target characters spanning different mesh structures and visual styles.

150
Evaluation animations

Total speech animations generated per method from 5 characters and 30 audio clips.

11.304
AnyTalk LSE-D

Lip Sync Error Distance; lower is better.

3.155
AnyTalk LSE-C

Lip Sync Error Confidence; higher is better.

110
AnyTalkRT speed

Real-time inference speed in frames per second, equivalent to 9.09 ms per frame.

20
CsF training time

Character-specific fine-tuning duration in minutes on an NVIDIA A6000.

What the paper found

AnyTalk converts audio into 3D speech animation for arbitrary rigged characters without paired audio-animation training data. It first adapts the Hallo audio-driven video diffusion model using Character-specific Fine-tuning, or CsF: rendered blendshape images are paired with zeroed audio embeddings, while temporal, audio-attention, and motion-attention modules remain frozen and only spatial residual layers adapt to the character. At inference, nonzero speech generates a character-consistent talking-head video with pose suppressed and lip motion emphasized. A second stage lifts this 2D motion into 3D by optimizing blendshape parameters against 14 talk-related landmarks, using expression-invariant landmark filtering, homography warping, asymmetric mouth-opening loss, and L1 regularization. Tests cover 5 characters with varied mesh structures and styles, using LibriSpeech audio; across 150 animations, AnyTalk achieves an LSE-D of 11.304 and LSE-C of 3.155, outperforming ScanTalk, DiffSpeaker plus NFR, and CodeTalker plus NFR. The method requires 20 minutes of CsF training on an NVIDIA A6000, but frame-wise optimization takes 3.12 seconds. To remove that bottleneck, AnyTalkRT distills the pipeline with feature matching and blendshape reconstruction losses, reaching 110 FPS, or 9.09 ms per frame, with some lip-sync degradation. The same CsF and landmark-optimization pipeline also transfers to MEMO, although the system remains dependent on frontal views and predefined blendshapes.

Original abstract

We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on character-specific training data or laborious rigging/re-meshing, AnyTalk circumvents these limitations by leveraging recent video diffusion models trained on extensive video datasets. We first adapt a pre-trained video diffusion model to a target character through our Character-specific Fine-tuning (\textit{CsF}) technique. By fine-tuning on rendered images of the 3D character paired with zeroed-out audio embeddings (representing "no motion"), we eliminate the need for animation data while preserving the motion prior of large-scale video diffusion model. We then uplift the resulting talking-head video into a 3D speech animation by estimating blendshape parameters through a proposed optimization process. AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements. We further enhance usability by distilling AnyTalk into a streamlined network, $\text{AnyTalk}_{RT}$, thereby enabling real-time performance. By leveraging talking-head video generation, our method broadens access to audio-driven speech animation technology for arbitrary characters. The code is publicly available at https://serin-yoon.github.io/projects/anytalk/.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis