CLJul 7

CoPiT: Cognitive Pivot Translation for Digraphic Low-Resource Mongolian in the Traditional Script

arXiv:2607.0584923.1Has Code
Predicted impact top 14% in CL · last 90 daysOriginality Incremental advance
AI Analysis

For researchers working on low-resource and digraphic languages, CoPiT provides a practical pivot-based method to leverage script resource hierarchies, though the approach is incremental.

CoPiT addresses the challenge of translating low-resource Mongolian in the Traditional script by routing translation through the Cyrillic script, achieving substantial BLEU improvements and 1.5-1.6x COMET gains, enabling open-source models to match or outperform GPT-4.1.

Low-resource languages remain challenging for machine translation, and Mongolian is a representative case. As a digraphic language, Mongolian is written in both Cyrillic and Traditional scripts, which exhibit a severe imbalance in data availability. While the Cyrillic script is relatively well-resourced, the Traditional script remains extremely data-scarce and orthographically ambiguous, leading to substantial performance degradation in direct translation. We propose CoPiT, a cognitively motivated pivot-based translation pipeline that exploits this internal resource hierarchy by routing translation through the Cyrillic script. The pipeline explicitly resolves script-induced ambiguity in the Traditional script before translation, enabling more stable and accurate meaning transfer. Across multiple backbone models and target languages, CoPiT consistently outperforms direct translation, achieving substantial absolute BLEU improvements together with consistent 1.5-1.6x COMET gains. These gains allow strong open-source models to match or outperform GPT-4.1 under comparable evaluation settings. Beyond inference-time improvements, CoPiT enables the construction of synthetic parallel data directly from Traditional-script text, mitigating data scarcity in realistic low-resource scenarios. We release a new multi-script parallel dataset covering Mongolian in both scripts alongside English, Korean, and Russian. All datasets and code are publicly available at https://anonymous.4open.science/r/anonymous_project-76C7.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes