Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis
AuthorsSanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu, David Kant, Dongmin Hyun, Dorian Desblancs, Gregory Antonovsky, Oleg Repin, Peng-Jen Chen, Xutai Ma, Zehai Tu, Juan Pino, Wei-Ning Hsu
Resources
Text-AB is a large-scale speech generation system that turns text into expressive dubbed voices and natural two-way conversations without requiring explicit text-speech alignment.
Key results
Main Text-AB model parameter count.
Hours of English and Spanish monolingual speech used for pretraining.
Frame rate of the latent representation for 48 kHz audio.
Word error rate after selecting among 32 generation candidates.
Points separating Text-AB from human recordings on the MOS scale.
Point gain over the prior internal dialogue synthesis system.
What the paper found
Meta’s Text-AB is an alignment-free speech generator for cross-lingual voice dubbing, stereo full-duplex dialogue, and emotion-controlled conversations. It combines a Diffusion Transformer with flow matching and DAC-VAE latent audio features, compressing 48 kHz waveforms into a 25 Hz sequence. Unlike Audiobox, it accepts raw text through the mT5 encoder and learns text–speech alignment with cross-attention, eliminating forced alignment and explicit duration prediction. The system scales to a 3B-parameter model pretrained on 480k hours of English and Spanish speech, then fine-tuned on dubbing and two-channel dialogue data. Multi-diffusion produces arbitrarily long conversations, while multi-stage reranking selects candidates using WER and WavLM speaker similarity; with 32 candidates, dubbing WER falls from 4.05% to 2.20% and speaker similarity rises from 0.66 to 0.74. On short-form dialogue, Text-AB comes within 0.09 points of human recordings in overall human-likeness, and on long-form synthetic scripts it improves human-likeness over the prior internal system by 0.86 points. Turn-level valence-arousal-dominance embeddings enable emotions to differ from the reference speaker’s tone. Evaluation uses Whisper-Large-V3, WavLM, Qwen2-Audio, and Gemini 2.5 Pro, showing stronger dubbing prosody, speaker preservation, naturalness, turn-taking, back-channeling, and emotional interaction than the baselines.
Original abstract
We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs from the Audiobox system along three dimensions. First, it operates in a latent diffusion framework using DAC-VAE features that encode 48 kHz waveforms into a 25 Hz latent sequence, giving over 10x higher compression than previous EnCodec representations while improving resynthesis quality. Second, Text-AB is alignment-free: it consumes raw text via an off-the-shelf text encoder and learns text-speech alignment through cross-attention, removing the need for forced alignment and explicit duration prediction. Third, we scale model and data substantially, pretraining a 3B-parameter model on 480k hours of monolingual speech, followed by supervised fine-tuning on three downstream tasks: cross-lingual voice dubbing, full-duplex dialogue synthesis, and emotional full-duplex dialogue synthesis. At inference, Text-AB supports one-shot generation for up to ~1 min of speech and arbitrarily long-form generation via a multi-diffusion scheme, plus a multi-stage reranking strategy that enhances quality based on automated metrics. On a real-world dubbing benchmark, Text-AB delivers a step-change improvement over the latest internal dubbing system, with large gains in prosody similarity, voice similarity, naturalness, and shareability. For full-duplex dialogue synthesis, it approaches human recordings on short-form conversations and substantially outperforms the latest internal model on long-form human-likeness and expressivity, while natively modeling turn-taking, back-channeling, and emotional dynamics. For emotional dialogue synthesis, emotion conditioning significantly improves emotion alignment and emotional interaction quality over the unconditioned baseline.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.