CLDec 12, 2018

SMT vs NMT: A Comparison over Hindi & Bengali Simple Sentences

Sainik Kumar Mahata, Soumil Mandal, Dipankar Das, Sivaji Bandyopadhyay

arXiv:1812.04898v11.320 citations

Originality Synthesis-oriented

AI Analysis

This work addresses translation quality for Hindi and Bengali speakers, but it is incremental as it compares established methods on specific data.

The study compared Statistical Machine Translation (SMT) and Neural Machine Translation (NMT) for English-Hindi and English-Bengali language pairs, finding that NMT outperforms SMT for simple sentences, while SMT performs better across all sentence types.

In the present article, we identified the qualitative differences between Statistical Machine Translation (SMT) and Neural Machine Translation (NMT) outputs. We have tried to answer two important questions: 1. Does NMT perform equivalently well with respect to SMT and 2. Does it add extra flavor in improving the quality of MT output by employing simple sentences as training units. In order to obtain insights, we have developed three core models viz., SMT model based on Moses toolkit, followed by character and word level NMT models. All of the systems use English-Hindi and English-Bengali language pairs containing simple sentences as well as sentences of other complexity. In order to preserve the translations semantics with respect to the target words of a sentence, we have employed soft-attention into our word level NMT model. We have further evaluated all the systems with respect to the scenarios where they succeed and fail. Finally, the quality of translation has been validated using BLEU and TER metrics along with manual parameters like fluency, adequacy etc. We observed that NMT outperforms SMT in case of simple sentences whereas SMT outperforms in case of all types of sentence.

View on arXiv PDF

Similar