A Survey of Methods to Leverage Monolingual Data in Low-resource Neural Machine Translation

1 Oct 2019 · Ilshat Gibadullin, Aidar Valeev, Albina Khusainova, Adil Khan ·

Neural machine translation has become the state-of-the-art for language pairs with large parallel corpora. However, the quality of machine translation for low-resource languages leaves much to be desired. There are several approaches to mitigate this problem, such as transfer learning, semi-supervised and unsupervised learning techniques. In this paper, we review the existing methods, where the main idea is to exploit the power of monolingual data, which, compared to parallel, is usually easier to obtain and significantly greater in amount.

PDF Abstract