Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition
This work addresses the practical deployment gap for ASR in 6 Southern Bantu languages spoken by over 80 million people, though results are incremental and show model-language interactions.
The authors developed a tone-conditioned curriculum learning framework for low-resource Southern Bantu speech recognition, achieving 28.41% average WER across datasets and 23.79% on Xitsonga transfer, but found no single model works best for all six languages.
Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and public services. We addressed this gap with a tone conditioned curriculum framework for 6 Southern Bantu languages that combined hybrid difficulty scoring, gated adapters driven by tonal statistics and staged curriculum training. We trained on a community corpus and tested transfer to NCHLT to measure robustness beyond matched evaluation. Results revealed clear interactions between architecture and language, with W2V-BERT outperforming Whisper on Nguni languages by 3 to 4 WER points whilst Whisper performed better on Sotho-Tswana languages. W2V-BERT with tone conditioning reached 28.41% average WER across datasets and 23.79% on Xitsonga transfer. No single model suited all 6 languages, so deployment should pair model selection per language with validation across corpora.