NTH

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation

AuthorsYexing Du, Kaiyuan Liu, Youcheng Pan, Bo Yang, Ming Liu, Bing Qin, Yang Xiang

June 2, 2026 2 min read
Watch on YouTube
The one-line take

This paper builds a privacy-preserving speech translation system that sends only compressed features from device to cloud, cutting bandwidth while scaling translation across 45 languages.

Key results

45
FLEURS languages

ESRT is evaluated on the FLEURS dataset across 45 languages for many-to-many speech-to-text translation.

44
FLEURS directions

The main many-to-many setup covers 45 × 44 translation directions on FLEURS.

83.8
ESRT-12B COMET X→44

ESRT-12B achieves the reported average COMET score for X→44 directions on FLEURS.

83.4
ESRT-12B COMET 44→X

ESRT-12B achieves the reported average COMET score for 44→X directions on FLEURS.

0.92
Raw WAV size

A 30-second mono WAV file is reported as about 0.92 MB before compression.

0.06
Tensor size

The compressed Q-Former tensor sent to the cloud is reported as about 0.06 MB.

What the paper found

Bandwidth-Efficient and Privacy-Preserving Edge-Cloud Many-to-Many Speech Translation introduces ESRT, a split-inference speech translation framework that keeps a frozen Whisper encoder and Q-Former-based speech adapter on-device, then sends only compressed intermediate tensors to a cloud LLM, instead of raw audio. This design cuts transmitted data from a 30-second WAV’s 0.92 MB to about 0.06 MB, a 15.6× reduction, and the paper reports up to 10× bandwidth savings in practical many-to-many usage. Privacy is strengthened through four mechanisms: lossy information bottlenecking, fixed-shape tensor transmission that obfuscates the backend model, 30-second padding that hides utterance duration, and implicit language encoding that avoids explicit language tags. To address English-centric bias and catastrophic forgetting, the authors add a multi-task weighted curriculum over ASR, speech-guided MT, and speech translation, with language-token vocabulary expansion and optional LoRA fine-tuning. On FLEURS, ESRT-12B reaches state-of-the-art COMET across 45 languages and 44 directions, averaging 83.8 for X→44 and 83.4 for 44→X, outperforming MCAT-Large-27B while using less than half the parameters; ESRT-4B also exceeds the 27B baseline on several settings. The paper’s broader claim is that strong many-to-many speech translation can be achieved without uploading raw voice, making edge-cloud deployment both more bandwidth-efficient and substantially safer for privacy.

Original abstract

Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment paradigms face critical challenges: pure on-device models suffer from resource constraints, while centralized cloud systems incur severe privacy risks and bandwidth bottlenecks by transmitting raw voice data. Furthermore, most models exhibit English-centric biases, restricting many-to-many translation scaling. In this paper, we propose Edge-cloud Speech Recognition and Translation (ESRT), a privacy-preserving and bandwidth-efficient collaborative edge-cloud MLLM framework. Specifically, we design an edge-cloud split inference architecture that retains a lightweight speech encoder and adapter on the device, transmitting only highly compressed intermediate features to the cloud. This fundamentally prevents voiceprint leakage and reduces bandwidth requirements by up to 10$\times$. To overcome English-centric bottlenecks, we introduce a multi-task weighted curriculum learning strategy with data balancing to ensure robust cross-lingual consistency. Extensive experiments on the FLEURS dataset demonstrate that our models, ESRT-4B and ESRT-12B, achieve state-of-the-art many-to-many S2TT performance across 45 languages ($45 \times 44$ directions). Code and models are released to facilitate reproducible, privacy-aware MLLM S2TT research. The code and models are released at https://github.com/yxduir/esrt.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis