NTH

Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

AuthorsSanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu, David Kant, Dongmin Hyun, Dorian Desblancs, Gregory Antonovsky, Oleg Repin, Peng-Jen Chen, Xutai Ma, Zehai Tu, Juan Pino, Wei-Ning Hsu

September 12, 2026 2 min read
Watch on YouTube
The one-line take

Text-AB is a large-scale speech generation system that turns text into expressive dubbed voices and natural two-way conversations without requiring explicit text-speech alignment.

Key results

3B
Model size

Main Text-AB model parameter count.

480k
Pretraining data

Hours of English and Spanish monolingual speech used for pretraining.

25 Hz
DAC-VAE latent rate

Frame rate of the latent representation for 48 kHz audio.

2.20%
Reranked dubbing WER

Word error rate after selecting among 32 generation candidates.

0.09
Short-form human-likeness gap

Points separating Text-AB from human recordings on the MOS scale.

0.86
Long-form human-likeness improvement

Point gain over the prior internal dialogue synthesis system.

What the paper found

Meta’s Text-AB is an alignment-free speech generator for cross-lingual voice dubbing, stereo full-duplex dialogue, and emotion-controlled conversations. It combines a Diffusion Transformer with flow matching and DAC-VAE latent audio features, compressing 48 kHz waveforms into a 25 Hz sequence. Unlike Audiobox, it accepts raw text through the mT5 encoder and learns text–speech alignment with cross-attention, eliminating forced alignment and explicit duration prediction. The system scales to a 3B-parameter model pretrained on 480k hours of English and Spanish speech, then fine-tuned on dubbing and two-channel dialogue data. Multi-diffusion produces arbitrarily long conversations, while multi-stage reranking selects candidates using WER and WavLM speaker similarity; with 32 candidates, dubbing WER falls from 4.05% to 2.20% and speaker similarity rises from 0.66 to 0.74. On short-form dialogue, Text-AB comes within 0.09 points of human recordings in overall human-likeness, and on long-form synthetic scripts it improves human-likeness over the prior internal system by 0.86 points. Turn-level valence-arousal-dominance embeddings enable emotions to differ from the reference speaker’s tone. Evaluation uses Whisper-Large-V3, WavLM, Qwen2-Audio, and Gemini 2.5 Pro, showing stronger dubbing prosody, speaker preservation, naturalness, turn-taking, back-channeling, and emotional interaction than the baselines.

Original abstract

We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs from the Audiobox system along three dimensions. First, it operates in a latent diffusion framework using DAC-VAE features that encode 48 kHz waveforms into a 25 Hz latent sequence, giving over 10x higher compression than previous EnCodec representations while improving resynthesis quality. Second, Text-AB is alignment-free: it consumes raw text via an off-the-shelf text encoder and learns text-speech alignment through cross-attention, removing the need for forced alignment and explicit duration prediction. Third, we scale model and data substantially, pretraining a 3B-parameter model on 480k hours of monolingual speech, followed by supervised fine-tuning on three downstream tasks: cross-lingual voice dubbing, full-duplex dialogue synthesis, and emotional full-duplex dialogue synthesis. At inference, Text-AB supports one-shot generation for up to ~1 min of speech and arbitrarily long-form generation via a multi-diffusion scheme, plus a multi-stage reranking strategy that enhances quality based on automated metrics. On a real-world dubbing benchmark, Text-AB delivers a step-change improvement over the latest internal dubbing system, with large gains in prosody similarity, voice similarity, naturalness, and shareability. For full-duplex dialogue synthesis, it approaches human recordings on short-form conversations and substantially outperforms the latest internal model on long-form human-likeness and expressivity, while natively modeling turn-taking, back-channeling, and emotional dynamics. For emotional dialogue synthesis, emotion conditioning significantly improves emotion alignment and emotional interaction quality over the unconditioned baseline.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis