NTH

GigaAM Multilingual: Foundation Model for Underrepresented Languages

AuthorsAndrei Kuzmenko, Alexandr Maximenko, Aleksandr Kutsakov, Georgii Gospodinov, Dmitrii Bolotov, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin

August 3, 2026 2 min read
Watch on YouTube
The one-line take

GigaAM Multilingual is a large speech foundation model and balancing recipe that improves ASR for data-poor Central Asian languages.

Key results

2M
Pre-training corpus

Hours of audio used to pre-train GigaAM Multilingual

600M
Main encoder size

Parameters in the 24-layer Conformer encoder

12.2%
Compact encoder average WER

Average WER of the 240M GigaAM encoder under matched CTC fine-tuning

3.8%
Georgian adaptation WER

Common Voice WER using separate-language CTC adaptation

3.6%
Bashkir adaptation WER

Common Voice WER using separate-language CTC adaptation

What the paper found

Researchers at SaluteDevices introduce GigaAM Multilingual, an open-weight Conformer speech-recognition foundation model targeting underrepresented Central Asian languages, especially Kazakh, Kyrgyz, and Uzbek. The encoder uses a HuBERT-style masked-unit prediction objective, Rotary Position Embeddings, and a 600M-parameter, 24-layer architecture trained on 2M hours of audio. Its main innovation is balancing multilingual data at the level of five language clusters during pre-training, then combining language- and domain-aware sampling during CTC fine-tuning to prevent English and Russian from dominating. The training mixture includes open-source, crowdsourced, synthetic, and weakly supervised speech, with synthetic data generated from mC4 text using an in-house TTS system with over 100 voices. On Common Voice, FLEURS, and internal spontaneous-speech tests, GigaAM substantially outperforms Whisper Large v3, Seamless-M4T v2 large, and Omnilingual ASR on the target languages. In matched CTC experiments, even the compact 240M GigaAM encoder reaches an average WER of 12.2%, versus 14.1% for Whisper Large v3 and 16.6% for Omnilingual-1B. For minimal-coverage adaptation using only Common Voice training data, GigaAM achieves 3.8% WER on Georgian and 3.6% on Bashkir, beating both baselines. The results show that structured sampling, rather than scale alone, can make multilingual ASR more effective and efficient for long-tail languages.

Original abstract

Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challenge of building robust foundation models for underrepresented Central Asian languages (Kazakh, Kyrgyz, Uzbek). We present GigaAM Multilingual, a Conformer encoder pre-trained on 2M hours of audio using a HuBERT-style objective. Crucially, we introduce a cluster-level data balancing strategy during pre-training and a domain-aware sampling method during fine-tuning to mitigate head-language dominance. In controlled comparisons, our approach outperforms strong open pretrained encoders (Whisper Large v3, Omnilingual-1B) on target languages, achieving significant gains on spontaneous speech while maintaining efficiency. We release the foundation encoder and ASR model, offering a proven recipe for effective multilingual adaptation under realistic data imbalance.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis