Papers
arxiv:2607.10371

GigaAM Multilingual: Foundation Model for Underrepresented Languages

Published on Jul 11
· Submitted by
Georgii Gospodinov
on Jul 21
Authors:
,
,
,
,
,

Abstract

Despite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challenge of building robust foundation models for underrepresented Central Asian languages (Kazakh, Kyrgyz, Uzbek). We present GigaAM Multilingual, a Conformer encoder pre-trained on 2M hours of audio using a HuBERT-style objective. Crucially, we introduce a cluster-level data balancing strategy during pre-training and a domain-aware sampling method during fine-tuning to mitigate head-language dominance. In controlled comparisons, our approach outperforms strong open pretrained encoders (Whisper Large v3, Omnilingual-1B) on target languages, achieving significant gains on spontaneous speech while maintaining efficiency. We release the foundation encoder and ASR model, offering a proven recipe for effective multilingual adaptation under realistic data imbalance.

Community

Paper author Paper submitter

GigaAM Multilingual — open-source (MIT) speech foundation models for underrepresented languages. Two Conformer encoders (220M / 600M) pre-trained HuBERT-style on 2M hours across 70+ languages, with charwise CTC decoders: best open-source WER on Kazakh, Kyrgyz, Uzbek — e.g. 32–86% relative WER reduction vs Whisper-large-v3 on Central Asian languages. The SSL backbones adapt to a new language with a few hours of labeled data

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Models citing this paper 6

Browse 6 models citing this paper

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2607.10371 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.