CL AIOct 11, 2021

Unsupervised Neural Machine Translation with Generative Language Models Only

Jesse Michael Han, Igor Babuschkin, Harrison Edwards, Arvind Neelakantan, Tao Xu, Stanislas Polu, Alex Ray, Pranav Shyam, Aditya Ramesh, Alec Radford, Ilya Sutskever

arXiv:2110.05448v14.943 citations

Originality Incremental advance

AI Analysis

This addresses the problem of translation without parallel data for NLP researchers, though it is incremental as it builds on existing pre-trained models.

The paper tackles unsupervised neural machine translation by leveraging generative language models, achieving a state-of-the-art BLEU score of 42.1 on the WMT14 English-French benchmark.

We show how to derive state-of-the-art unsupervised neural machine translation systems from generatively pre-trained language models. Our method consists of three steps: few-shot amplification, distillation, and backtranslation. We first use the zero-shot translation ability of large pre-trained language models to generate translations for a small set of unlabeled sentences. We then amplify these zero-shot translations by using them as few-shot demonstrations for sampling a larger synthetic dataset. This dataset is distilled by discarding the few-shot demonstrations and then fine-tuning. During backtranslation, we repeatedly generate translations for a set of inputs and then fine-tune a single language model on both directions of the translation task at once, ensuring cycle-consistency by swapping the roles of gold monotext and generated translations when fine-tuning. By using our method to leverage GPT-3's zero-shot translation capability, we achieve a new state-of-the-art in unsupervised translation on the WMT14 English-French benchmark, attaining a BLEU score of 42.1.

View on arXiv PDF

Similar