CLJun 6, 2022

MorisienMT: A Dataset for Mauritian Creole Machine Translation

arXiv:2206.02421v10.61 citations
Originality Synthesis-oriented
AI Analysis

This provides a benchmark for machine translation in Mauritian Creole, addressing a gap for a widely spoken but under-resourced language, though it is incremental as it focuses on dataset creation.

The authors tackled the lack of resources for machine translation of Mauritian Creole by creating MorisienMT, a dataset including parallel corpora with English and French, and established baseline models using this data and transfer learning.

In this paper, we describe MorisienMT, a dataset for benchmarking machine translation quality of Mauritian Creole. Mauritian Creole (Morisien) is the lingua franca of the Republic of Mauritius and is a French-based creole language. MorisienMT consists of a parallel corpus between English and Morisien, French and Morisien and a monolingual corpus for Morisien. We first give an overview of Morisien and then describe the steps taken to create the corpora and, from it, the training and evaluation splits. Thereafter, we establish a variety of baseline models using the created parallel corpora as well as large French--English corpora for transfer learning. We release our datasets publicly for research purposes and hope that this spurs research for Morisien machine translation.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes