GigaAM Multilingual: Foundation Model for Underrepresented Languages
AuthorsAndrei Kuzmenko, Alexandr Maximenko, Aleksandr Kutsakov, Georgii Gospodinov, Dmitrii Bolotov, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin
Resources
GigaAM Multilingual is a large speech foundation model and balancing recipe that improves ASR for data-poor Central Asian languages.
Key results
Hours of audio used to pre-train GigaAM Multilingual
Parameters in the 24-layer Conformer encoder
Average WER of the 240M GigaAM encoder under matched CTC fine-tuning
Common Voice WER using separate-language CTC adaptation
Common Voice WER using separate-language CTC adaptation
What the paper found
Researchers at SaluteDevices introduce GigaAM Multilingual, an open-weight Conformer speech-recognition foundation model targeting underrepresented Central Asian languages, especially Kazakh, Kyrgyz, and Uzbek. The encoder uses a HuBERT-style masked-unit prediction objective, Rotary Position Embeddings, and a 600M-parameter, 24-layer architecture trained on 2M hours of audio. Its main innovation is balancing multilingual data at the level of five language clusters during pre-training, then combining language- and domain-aware sampling during CTC fine-tuning to prevent English and Russian from dominating. The training mixture includes open-source, crowdsourced, synthetic, and weakly supervised speech, with synthetic data generated from mC4 text using an in-house TTS system with over 100 voices. On Common Voice, FLEURS, and internal spontaneous-speech tests, GigaAM substantially outperforms Whisper Large v3, Seamless-M4T v2 large, and Omnilingual ASR on the target languages. In matched CTC experiments, even the compact 240M GigaAM encoder reaches an average WER of 12.2%, versus 14.1% for Whisper Large v3 and 16.6% for Omnilingual-1B. For minimal-coverage adaptation using only Common Voice training data, GigaAM achieves 3.8% WER on Georgian and 3.6% on Bashkir, beating both baselines. The results show that structured sampling, rather than scale alone, can make multilingual ASR more effective and efficient for long-tail languages.
Original abstract
Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challenge of building robust foundation models for underrepresented Central Asian languages (Kazakh, Kyrgyz, Uzbek). We present GigaAM Multilingual, a Conformer encoder pre-trained on 2M hours of audio using a HuBERT-style objective. Crucially, we introduce a cluster-level data balancing strategy during pre-training and a domain-aware sampling method during fine-tuning to mitigate head-language dominance. In controlled comparisons, our approach outperforms strong open pretrained encoders (Whisper Large v3, Omnilingual-1B) on target languages, achieving significant gains on spontaneous speech while maintaining efficiency. We release the foundation encoder and ASR model, offering a proven recipe for effective multilingual adaptation under realistic data imbalance.
Read the original paperMore in Speech AI
Browse all 27 papers →Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen
A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.
Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry
Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee
The study argues that voice-similarity systems should be judged by whether their embedding geometry matches human perception, not merely by verification accuracy.
Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
Jian Chen, You Zhang, Mark Vinton
Tacit-TTS makes zero-shot voice cloning over ten times faster while preserving the ability to clone voices from speech without transcripts, including multilingual and non-lexical references.