HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
AuthorsBohan Li, Shi Lian, Hankun Wang, Yiwei Guo, Yu Xi, Zhihan Li, Da Zheng, Colin Zhang, Kai Yu
HoliTok is a new speech tokenization approach that compresses audio into a compact latent sequence usable for both speech generation and understanding, helping unify speech AI systems.
Key results
HoliTok reconstruction on LibriSpeech test-other
HoliTok uses a compact latent representation for reconstruction
HoliTok encodes speech at 25 Hz
HoliTok-Unite average ASR WER over HoliTok-Base in unified spoken language modeling
What the paper found
HoliTok, from Shanghai Jiao Tong University’s X-LANCE Lab with hi lab at Xiaohongshu Inc., proposes a continuous speech tokenizer for unified generation-understanding systems, targeting a representation that is simultaneously decodable, compact, and learnable by a language model. It encodes 48 kHz audio into a 25 Hz sequence of 128-dimensional latents using a low-latency VAE with causal strided convolutions, a temporal variational bottleneck, and a BigVGAN-style decoder. Its key novelty is a three-stage progressive training recipe: first a deterministic autoencoder for waveform fidelity, then weak KL regularization to convert that manifold into a smooth stochastic latent space, and finally downstream-aware enrichment with WavLM frame-level distillation, x-vector utterance-level distillation, and task-conditioned supervision for ASR, emotion recognition, audio captioning, and sound event detection. In unified AR+DiT experiments using a Qwen2.5-0.5B backbone, HoliTok is the only representation that works robustly without extra optimization tricks. On LibriSpeech test-other, it reaches PESQ 4.10/4.01, STOI 0.974, WER 4.22%, SPKSIM 0.968, and EMOSIM 0.995 at a 7.5× compression ratio and 25 tokens per second. In zero-shot TTS, it outperforms Semantic-VAE and MingTok-Audio on Seed-TTS-Eval and Emergent-TTS, and the HoliTok-Unite variant improves unified ASR–TTS balance, cutting average TTS WER from 20.90% to 8.59% and ASR WER from 12.63% to 8.02% over HoliTok-Base.
Original abstract
Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenizers, however, often fail to satisfy these requirements simultaneously, leading to increased architectural complexity and more involved training designs. We propose HoliTok, a continuous Holistic speech Tokenization model designed for unified generation-understanding modeling. HoliTok encodes 48~kHz speech into a compact 25~Hz sequence of 128-dimensional latents. It is trained with a progressive strategy that jointly preserves signal-level fidelity, incorporates semantic information, and maintains strong latent learnability. Based on this tokenization, we build a unified AR+DiT model for speech synthesis and recognition, where the same latent sequence supports both generation-specific and unified generation-understanding tasks. Experiments show that HoliTok achieves competitive reconstruction fidelity, improves generative learnability for high-quality and controllable synthesis, and, among the evaluated representations, is the only one that operates robustly in our unified generation-understanding architecture without additional optimization tricks. These results suggest that HoliTok serves as an effective speech tokenizer and a foundational representation interface for unified spoken language modeling. The code is available at: https://github.com/bovod-sjtu/HoliTok.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.