Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation
AuthorsYexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Ming Liu, Bing Qin, Yang Xiang
This paper builds a privacy-preserving speech translation system that sends only compressed features from device to cloud, cutting bandwidth while scaling translation across 45 languages.
Key results
ESRT is evaluated on the FLEURS dataset across 45 languages for many-to-many speech-to-text translation.
The main many-to-many setup covers 45 × 44 translation directions on FLEURS.
ESRT-12B achieves the reported average COMET score for X→44 directions on FLEURS.
ESRT-12B achieves the reported average COMET score for 44→X directions on FLEURS.
A 30-second mono WAV file is reported as about 0.92 MB before compression.
The compressed Q-Former tensor sent to the cloud is reported as about 0.06 MB.
What the paper found
Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation introduces ESRT, a split-inference speech translation framework that keeps a frozen Whisper encoder and Q-Former-based speech adapter on-device, then sends only compressed intermediate tensors to a cloud LLM, instead of raw audio. This design cuts transmitted data from a 30-second WAV’s 0.92 MB to about 0.06 MB, a 15.6× reduction, and the paper reports up to 10× bandwidth savings in practical many-to-many usage. Privacy is strengthened through four mechanisms: lossy information bottlenecking, fixed-shape tensor transmission that obfuscates the backend model, 30-second padding that hides utterance duration, and implicit language encoding that avoids explicit language tags. To address English-centric bias and catastrophic forgetting, the authors add a multi-task weighted curriculum over ASR, speech-guided MT, and speech translation, with language-token vocabulary expansion and optional LoRA fine-tuning. On FLEURS, ESRT-12B reaches state-of-the-art COMET across 45 languages and 44 directions, averaging 83.8 for X→44 and 83.4 for 44→X, outperforming MCAT-Large-27B while using less than half the parameters; ESRT-4B also exceeds the 27B baseline on several settings. The paper’s broader claim is that strong many-to-many speech translation can be achieved without uploading raw voice, making edge-cloud deployment both more bandwidth-efficient and substantially safer for privacy.
Original abstract
Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment paradigms face critical challenges: pure on-device models suffer from resource constraints, while centralized cloud systems incur severe privacy risks and bandwidth bottlenecks by transmitting raw voice data. Furthermore, most models exhibit English-centric biases, restricting many-to-many translation scaling. In this paper, we propose Edge-cloud Speech Recognition and Translation (ESRT), a privacy-preserving and bandwidth-efficient collaborative edge-cloud MLLM framework. Specifically, we design an edge-cloud split inference architecture that retains a lightweight speech encoder and adapter on the device, transmitting only highly compressed intermediate features to the cloud. This fundamentally prevents voiceprint leakage and reduces bandwidth requirements by up to 10$\times$. To overcome English-centric bottlenecks, we introduce a multi-task weighted curriculum learning strategy with data balancing to ensure robust cross-lingual consistency. Extensive experiments on the FLEURS dataset demonstrate that our models, ESRT-4B and ESRT-12B, achieve state-of-the-art many-to-many S2TT performance across 45 languages ($45 \times 44$ directions). Code and models are released to facilitate reproducible, privacy-aware MLLM S2TT research. The code and models are released at https://github.com/yxduir/esrt.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.