A Shocking Amount of the Web is Machine Translated: Insights from Multi-Way Parallelism
Brian Thompson, Mehak Preet Dhaliwal, Peter Frisch, Tobias Domhan, Marcello Federico
Introduction
Modern AI is enabled by huge amounts of training data, typically several hundred billion tokens to a few trillion tokens Sun et al. (2021); Chowdhery et al. (2023); Touvron et al. (2023); Almazrouei et al. (2023). Training at this scale is only possible with web-scraped data.
We explore the effects that the long-term availability of low cost Machine Translation (MT) has had on the web.Free MT has been available online since late 1997 Gaspari and Hutchins (2007), around the same time that MT researchers began scraping the web for training data Resnik (1998). We show that content on the web is often translated into many languages, and the quality of these multi-way translations indicates they were primarily created using MT: see Figure 1. Machine generated, multi-way parallel translations not only dominate the total amount of translated content on the web in lower resource languages, it also constitutes a large fraction of the total web content in those languages. We also find evidence of a selection bias in the type of content which is translated into many languages, and therefore over represented in lower resource languages: This content is shorter, more predictable, and has a different topic distribution compared to content translated into a single language. A limited investigation suggests this selection bias is the result of low quality content generated in English (likely produced to generate ad revenue) and translated en masse into many lower resource languages via MT (again, likely to generate ad revenue).
Our findings raise numerous concerns for multilingual model builders: Fluency (especially across sentences) and accuracy are lower for MT data,MT technology has improved dramatically over the last decade, but still falls short of human quality Freitag et al. (2023). MT content has been added to the web over many years using MT systems available at the time, so much of the MT on the web is likely very low quality by modern standards. which could produce less fluent models with more hallucinations, and the selection bias indicates the data may be of lower quality, even before considering MT errors. Data quality is crucial in Large Language Model (LLM) training, where high quality corpora like books and Wikipedia articles are typically upsampled several times Brown et al. (2020); Gao et al. (2020); Rae et al. (2021); Le Scao et al. (2022).
Our findings also help to explain why low-resource MT Khan et al. (2017); Duh (2018); NLLB Team et al. (2022) is challenging, and why filtering noise Khayrallah and Koehn (2018) from web-scraped bitext Junczys-Dowmunt (2018); Chaudhary et al. (2019) is beneficial for MT training Koehn et al. (2018, 2019, 2020); Sloto et al. (2023).
To enable analysis, we create the largest multi-way corpus to date, consisting of 6.4B unique sentences in 90 languages. We release code to reproduce our corpus and analysis.https://github.com/amazon-science/multi-way-parallel-ccmatrix/. Corpus creation has been optimized to run in about one day on a single i4i.32xlarge AWS instance.
Related Work
Our work is inspired by several recent efforts which seek to understand the characteristics of large scale corpora Mehmood et al. (2017); Dodge et al. (2021); Kreutzer et al. (2022); Brannon et al. (2023). Many works have detected machine translation Kurokawa et al. (2009); Arase and Zhou (2013); Aharoni et al. (2014), but we are not aware of prior work using multi-way parallelism to do so. Freitag and Firat (2020) explored multi-way parallelism with the goal of improving multilingual MT.
Exploring multi-way parallelism on the web requires a curated representation of translated content from the web. We build upon ccMatrix Schwenk et al. (2021), which is in turn based on Common Crawl.https://commoncrawl.org/ Common Crawl is a long running web-scraping project which maintains a free, open source repository of web-scraped data. ccMatrix is created by embedding Common Crawl sentences into a multilingual space using LASER Artetxe and Schwenk (2019) and then finding bilingual translation pairs using fast approximate nearest neighbor search Johnson et al. (2019). We choose ccMatrix over a corpus from a traditional bitext mining process of document alignment Resnik and Smith (2003); Buck and Koehn (2016); Thompson and Koehn (2020) followed by sentence alignment Gale and Church (1993); Sennrich and Volk (2010); Thompson and Koehn (2019), for several reasons: is is the largest corpus available at the time of writing, sentence pairs have associated LASER margin scores, and we expect it to be a more general representation of the web than a corpus like Paracrawl, which intentionally targets bitext-rich domains Bañón et al. (2020).
Corpus Creation: MWccMatrix
We create a multi-way parallel representation of the web, consisting of translation tuples containing two or more sentences in different languages which are translations of each other.Unless otherwise noted, we use the term “translation” to mean a sentence which appears in a translation tuple – i.e. we do not attempt to distinguish whether that sentence was translated into or out of a given language. As a trivial example, (“hello”, “hola”) in English-Spanish and (“hello”, “olá”) in English-Portuguese combine to make (En:“hello”, Es:“hola”, Pt:“olá”). We denote this corpus Multi-Way ccMatrix (MWccMatrix).
We iterate through all bitext in ccMatrix, from highest to lowest LASER margin score, adding sentence pairs as new tuples in MWccMatrix when neither sentence is already in the new corpus, and expanding tuples already in the new corpus when one sentence or the other (but not both) is already present. This deduplicates the corpus (i.e. adds each unique sentence only once), but allows for more than one sentence in the same language to be added to a given tuple, which tend to differ primarily in punctuation/capitalization (i.e. near duplicates). Therefore, we remove all but the first sentence added to each tuple in a given language. Deduplication across language pairs brings the total number of sentences down from 21.7B total sentences (10.9B sentence pairs) to 7.9B unique sentences in 2.2B tuples, and near duplicate removal brings it down to 6.4B. Pseudocode and a description of the optimizations required to make corpus creation tractable are provided in Appendix A.
Analysis
We compared the total number of unique sentences (before removing near-duplicates) in MWccMatrix to the total number of unique sentences from the Common Crawl snapshots that the data is based on, as reported by Schwenk et al. (2021). They only report the number of unique sentences for the 54 (of 90) largest resource languages, so we cannot compute the fraction of sentences with one or more translations in the 36 lowest-resource languages.Measuring the number of unique sentences in the lowest resource languages would require re-processing the Common Crawl snapshots, which is computationally prohibitive. The percentage of unique monolingual sentences which have at least one translation is quite high, even for some high resource languages (e.g. 9.4% of English, 17.5% of French): see Figure 2.
2 Translations on the Web are Highly Multi-way Parallel
Of the 6.38B sentences in our 2.19B translation tuples, 3.63B (57.1%) are in multi-way parallelWe use “multi-way parallelism” (or simply “parallelism”) to refer to the size of the translation tuple that that sentence is in. For example, a sentence with parallelism of 5 comes from a tuple of size 5, which contains the given sentence plus translations in 4 other languages. (3+ languages) tuples: see Table 1. lower resource languages tend to be more multi-way parallel, with the 10 highest-resourced languages in ccMatrix having an average parallelism of 4.0, and the 10 lowest-resource languages in ccMatrix having an average parallelism of 8.6: see Figure 3.
3 Multi-way Parallel Translations are Lower Quality
We evaluate the quality of translations on the web using Quality Estimation (QE), with the CometQE model Rei et al. (2022), across different levels of multi-way parallelism. Modern quality estimation methods are nearly on par with reference-based metrics Freitag et al. (2023) and have been shown to perform well on noisy web data Peter et al. (2023). As QE does not require human annotation or human references, it allows us to evaluate a very large data sample (1M samples per language pair) and many language pairs.We select from WMT language pairs as CometQE is trained on WMT metrics annotations, thus we expect CometQE to be most accurate in those language pairs.
We find that highly multi-way parallel translations are significantly lower quality (6.2 CometQE points worse) than 2-way parallel translations. This trend is consistent across all 8 language pair directions we considered: see Table 2
4 Multi-way Parallel Data has Different Topic Distribution
We observe that multi-way parallel data consists of shorter, more predictable sentences: see Appendix D. To better understand this finding, we hired professional linguists to classify 10,000 randomly selected English sentences as one of the 20 topics given in Table 3, based loosely on the high-level Topics API categories.https://cloud.google.com/natural-language/docs/categories We observe a fairly dramatic shift in the distribution of topics when comparing 2-way parallel data to 8+ way parallel data, with Conversation & Opinion increasing from 22.5% to 40.1%.
We manually inspected a random sample of 100 highly multi-way parallel sentences from the Conversation & Opinion topic and found them hard to characterize due to the isolated sentences being very short (typically 5-10 words). However, searching the web for the sentences was enlightening: the vast majority came from articles that we characterized as low quality, requiring little or no expertise or advance effort to create, on topics like being taken more seriously at work, being careful about your choices, six tips for new boat owners, deciding to be happy, etc. Furthermore, we were unable to find any translationese or other errors that would suggest the articles were being translated into English (either by human translators or MT), suggesting it is instead being generated in English and translated to other languages.
Discussion & Conclusion
Experiments with both QE (§ 4.3) and LASER (see Appendix C) strongly suggest that highly multi-way parallel translations are generated by MT. In lower resource languages, most translations are multi-way parallel (§ 4.2), suggesting that MT content dominates translation content. Furthermore, a large fraction of the total sentences in lower resource languages have at least one translation (§ 4.1), implying that a large fraction of the total web in those languages is MT generated.
Several observations point to a selection bias in the type of data which is translated into many languages, compared to data translated into a single language: it is shorter and more predictable (Appendix D), and substantially more likely to be from the Conversation & Opinion topic (§ 4.4). Since translations of this data constitute a substantial portion of the total data in low-resource languages, this bias will also appear in low resource languages. Investigation of the increase in Conversation & Opinion data by the authors suggest that this selection bias is the result of low quality content (likely produced to generate ad revenue) being translated via MT en masse into many lower resource languages (again likely for the purpose of generating ad revenue). It also suggests that such data originates in English and is translated into other languages. Additional investigation would be required to understand if this finding generalizes to other topics, languages, and levels of multi-way parallelism.
Our findings also point to some ways to address the problem of MT output in web-scraped training data: It suggests that MT detection, which has typically been proposed to filter bitext, could also be helpful in filtering monolingual text in lower resource languages. It also suggests that multi-way parallelism is a promising way to detect low quality, machine translated data, especially in lower resource languages, to filter both bilingual and monolingual data.
References
Appendix A MWccMatrix Creation: Additional Details
A simplified version of the algorithm used to create MWccMatrix is provided in algorithm 1.
In practice, several optimizations were required to make the process tractable. Instead of attempting to sort 10.9B sentence pairs by margin score, we approximate the search by binning margin scores and sorting the data into buckets corresponding to the (binned) margin scores, similar to a radix sort. The sentences are too large to fit in memory, so we represent the sentences as 64 bit hashes. Additionally, our scripts are written in python but we use the cykhashhttps://github.com/realead/cykhash package, which provides a native C int64 to int64 hashmap. Conversion from hashes back to sentences is done in small shards, and the hashsent mappings required to reconstruct the data are sharded such that only the mappings required for one shard are loaded in memory at one time. Finally, we make extensive use of parallelization (e.g. computing margin score bins, sharding data by margin score bin, hashing sentence pairs, etc).
Appendix B Larger Version of Figure 2
A larger version of Figure 2, which includes language codes for each language, is provided in Figure 4. As previously noted, we only have total data sizes for the 54 highest-resource languages, as that is what was reported by Schwenk et al. (2021), so we cannot compute this percentage for the 36 lowest-resource languages used in this study.
Appendix C LASER Margin Score Analysis
We investigate how LASER margin score varies with multi-way parallelism. We find that multi-way parallel data tends to have higher margin scores: see Figure 5. Further investigation reveals that LASER has a strong bias for MT output over human translations (see Table 4), thus LASER margin scores for more multi-way parallel content are consistent with multi-way parallel data being MT. LASER’s preference for MT is likely because LASER is based on a small MT model. Similar phenomenon has been observed Freitag et al. (2021) in the Prism metric Thompson and Post (2020a, b), which is also based on an MT model.
Appendix D Length & Perplexity Analysis
We perform monolingual analysis to explore how data varies with multi-way parallelism. We find that more multi-way parallel sentences are shorter in length (see Table 5) and have lower perplexity (i.e. are easier to predict) as measured by GPT-2 Radford et al. (2019): see Figure 6. Since length could interact with cometQE scores, we verified that the results of § 4.3 also hold across sentence length: see Figure 7.