GigaAM Multilingual: Foundation Model for Underrepresented Languages
Abstract
Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challenge of building robust foundation models for underrepresented Central Asian languages (Kazakh, Kyrgyz, Uzbek). We present GigaAM Multilingual, a Conformer encoder pre-trained on 2M hours of audio using a HuBERT-style objective. Crucially, we introduce a cluster-level data balancing strategy during pre-training and a domain-aware sampling method during fine-tuning to mitigate head-language dominance. In controlled comparisons, our approach outperforms strong open pretrained encoders (Whisper Large v3, Omnilingual-1B) on target languages, achieving significant gains on spontaneous speech while maintaining efficiency. We release the foundation encoder and ASR model, offering a proven recipe for effective multilingual adaptation under realistic data imbalance.
Community
GigaAM Multilingual — open-source (MIT) speech foundation models for underrepresented languages. Two Conformer encoders (220M / 600M) pre-trained HuBERT-style on 2M hours across 70+ languages, with charwise CTC decoders: best open-source WER on Kazakh, Kyrgyz, Uzbek — e.g. 32–86% relative WER reduction vs Whisper-large-v3 on Central Asian languages. The SSL backbones adapt to a new language with a few hours of labeled data
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- From Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech Recognition (2026)
- Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings (2026)
- UR-BERT: Scaling Text Encoders for Massively Multilingual TTS Through Universal Romanization and Speech Token Prediction (2026)
- Building Community-Centred NLP Resources for Puno Quechua (2026)
- QuaSR: Quality-Aware Sample Reweighting for Pacific Indigenous Speech Recognition (2026)
- From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages (2026)
- Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 6
ai-babai/gigaam-multilingual-mlx
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper