Language and translation model adaptation using comparable corpora

Title	Language and translation model adaptation using comparable corpora
Publication Type	Conference Papers
Year of Publication	2008
Authors	Snover M, Dorr BJ, Schwartz R
Conference Name	Proceedings of the Conference on Empirical Methods in Natural Language Processing
Date Published	2008///
Publisher	Association for Computational Linguistics
Conference Location	Stroudsburg, PA, USA
Abstract	Traditionally, statistical machine translation systems have relied on parallel bi-lingual data to train a translation model. While bi-lingual parallel data are expensive to generate, monolingual data are relatively common. Yet monolingual data have been under-utilized, having been used primarily for training a language model in the target language. This paper describes a novel method for utilizing monolingual target data to improve the performance of a statistical machine translation system on news stories. The method exploits the existence of comparable text---multiple texts in the target language that discuss the same or similar stories as found in the source language document. For every source document that is to be translated, a large monolingual data set in the target language is searched for documents that might be comparable to the source documents. These documents are then used to adapt the MT system to increase the probability of generating texts that resemble the comparable document. Experimental results obtained by adapting both the language and translation models show substantial gains over the baseline system.
URL	http://dl.acm.org/citation.cfm?id=1613715.1613825

Language and translation model adaptation using comparable corpora

Publications