CLJun 24

Riazi-8B: An Urdu Large Language Model for Mathematical Reasoning

arXiv:2606.2556816.3
Predicted impact top 59% in CL · last 90 daysOriginality Synthesis-oriented
AI Analysis

For Urdu-speaking users, this work addresses the lack of reasoning-oriented resources and models in Urdu, enabling multi-step mathematical problem solving in a low-resource language.

Riazi-8B, an Urdu LLM for mathematical reasoning, achieves consistent improvements in answer correctness, reasoning quality, response completeness, and Urdu generation over existing Urdu instruction-tuned models on MGSM-Urdu, through continued pre-training on Urdu Wikipedia and supervised fine-tuning on Urdu Chain-of-Thought data from GSM8K.

Recent LLMs demonstrate strong mathematical reasoning capabilities, but existing gains rely heavily on English-centric training resources and benchmarks. As a result, reasoning performance degrades substantially in low-resource languages such as Urdu, where reasoning-oriented datasets and adapted models remain scarce. Urdu lacks both reasoning-oriented resources and models adapted for multi-step mathematical problem solving, limiting the applicability of recent progress to Urdu-speaking users. We address this gap through Riazi-8B, an Urdu mathematical reasoning model developed through a two-step adaptation process comprising continued pre-training on Urdu Wikipedia and supervised fine-tuning on Urdu Chain-of-Thought data derived from GSM8K. We evaluate Riazi-8B on MGSM-Urdu against existing Urdu instruction-tuned models. Our results show consistent improvements in answer correctness, reasoning quality, response completeness, and Urdu generation. Our findings demonstrate that combining Urdu language adaptation with reasoning-focused fine-tuning is an effective strategy for extending mathematical reasoning capabilities to low-resource languages.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes