NTH

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

AuthorsBohan Li, Shi Lian, Hankun Wang, Yiwei Guo, Yu Xi, Zhihan Li, Da Zheng, Colin Zhang, Kai Yu

June 2, 2026 2 min read
Watch on YouTube
The one-line take

HoliTok is a new speech tokenization approach that compresses audio into a compact latent sequence usable for both speech generation and understanding, helping unify speech AI systems.

Key results

4.22%
LibriSpeech WER

HoliTok reconstruction on LibriSpeech test-other

7.5x
Compression ratio

HoliTok uses a compact latent representation for reconstruction

25
Tokens per second

HoliTok encodes speech at 25 Hz

8.02%
Unified ASR WER

HoliTok-Unite average ASR WER over HoliTok-Base in unified spoken language modeling

What the paper found

HoliTok, from Shanghai Jiao Tong University’s X-LANCE Lab with hi lab at Xiaohongshu Inc., proposes a continuous speech tokenizer for unified generation-understanding systems, targeting a representation that is simultaneously decodable, compact, and learnable by a language model. It encodes 48 kHz audio into a 25 Hz sequence of 128-dimensional latents using a low-latency VAE with causal strided convolutions, a temporal variational bottleneck, and a BigVGAN-style decoder. Its key novelty is a three-stage progressive training recipe: first a deterministic autoencoder for waveform fidelity, then weak KL regularization to convert that manifold into a smooth stochastic latent space, and finally downstream-aware enrichment with WavLM frame-level distillation, x-vector utterance-level distillation, and task-conditioned supervision for ASR, emotion recognition, audio captioning, and sound event detection. In unified AR+DiT experiments using a Qwen2.5-0.5B backbone, HoliTok is the only representation that works robustly without extra optimization tricks. On LibriSpeech test-other, it reaches PESQ 4.10/4.01, STOI 0.974, WER 4.22%, SPKSIM 0.968, and EMOSIM 0.995 at a 7.5× compression ratio and 25 tokens per second. In zero-shot TTS, it outperforms Semantic-VAE and MingTok-Audio on Seed-TTS-Eval and Emergent-TTS, and the HoliTok-Unite variant improves unified ASR–TTS balance, cutting average TTS WER from 20.90% to 8.59% and ASR WER from 12.63% to 8.02% over HoliTok-Base.

Original abstract

Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenizers, however, often fail to satisfy these requirements simultaneously, leading to increased architectural complexity and more involved training designs. We propose HoliTok, a continuous Holistic speech Tokenization model designed for unified generation-understanding modeling. HoliTok encodes 48~kHz speech into a compact 25~Hz sequence of 128-dimensional latents. It is trained with a progressive strategy that jointly preserves signal-level fidelity, incorporates semantic information, and maintains strong latent learnability. Based on this tokenization, we build a unified AR+DiT model for speech synthesis and recognition, where the same latent sequence supports both generation-specific and unified generation-understanding tasks. Experiments show that HoliTok achieves competitive reconstruction fidelity, improves generative learnability for high-quality and controllable synthesis, and, among the evaluated representations, is the only one that operates robustly in our unified generation-understanding architecture without additional optimization tricks. These results suggest that HoliTok serves as an effective speech tokenizer and a foundational representation interface for unified spoken language modeling. The code is available at: https://github.com/bovod-sjtu/HoliTok.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis