NTH

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

AuthorsYifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

AffiliationsShanghai Jiao Tong University · University of Cambridge · Tsinghua University · Alibaba Token Hub, Alibaba Group · Nankai University · Chinese University of Hong Kong · SII https://github.com/yfyeung/PrunedCTC

October 2, 2026 2 min read
Watch on YouTube
The one-line take

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Key results

5.1×
Full-step memory reduction

Reduction for Zipformer-M with a 180K vocabulary

17%
Training step overhead

Additional time accompanying the 5.1× memory reduction

10×
LLM-CTC recognition speed

Maximum speedup over autoregressive LLM-CE on GigaSpeech

3%
Streaming relative WER gap

Maximum gap from matched offline Qwen3-ASR models on GigaSpeech test

What the paper found

Pruned CTC makes native large-vocabulary CTC practical for ASR by separating alignment computation from normalization: because valid CTC paths contain only transcript tokens and blank, it computes the dynamic program over that small subset while preserving the full-vocabulary softmax normalizer, yielding exactly the same loss and first-order gradients as standard CTC. Chunked vocabulary projection and finite-beam alignment pruning further reduce activation storage; with a Zipformer-M encoder, a 180K vocabulary, and an NVIDIA H100, full-step memory drops 5.1× with only 17% step-time overhead, while matching standard CTC on LibriSpeech, GigaSpeech, and AISHELL-1. Building on this method, LLM-CTC adapts causal Qwen3 models from 0.6B to 32B parameters for non-autoregressive recognition over their native vocabularies, avoiding sequential decoding. On GigaSpeech, it stays within 7% relative WER of autoregressive LLM-CE while achieving up to 10× faster recognition. A bounded-history streaming version reuses the KV cache and trains from utterance-level transcripts without chunk-level alignments; fine-tuned Qwen3-ASR models remain within 3% relative WER of matched offline models.

Original abstract

Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1$\times$ with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10$\times$ faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.

Read the original paper

More in Speech AI

Browse all 27 papers →
03Speech

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong, Phatrasek Jirabovonvisut

A large voice-cloning model becomes a synthetic data generator for training a compact, reference-free Thai TTS system that runs on-device.

Read analysis