Revisiting non-English Text Simplification: A Unified Multilingual Benchmark
Michael J. Ryan, Tarek Naous, Wei Xu
Introduction
Automatic text simplification (ATS) is the task of reducing the complexity of a text without changing its original content and meaning (Al-Thanyyan and Azmi, 2021). ATS has many applications, from making a text easier to read for people with reading and cognitive disabilities (Stajner, 2021) and second language learners (Petersen and Ostendorf, 2007) to reducing the complexity of medical texts for easier understanding by the general public (van den Bercken et al., 2019). For better accessibility to diverse communities, this technology should be available without language barriers.
Much of the recent success in English text simplification comes from large parallel corpora of texts with the same content written using both complicated and simple sentences (Xu et al., 2015; Jiang et al., 2020; Alva-Manchego et al., 2020). These resources enable the training of large language models for ATS in English (Scarton and Specia, 2018; Martin et al., 2020; Omelianchuk et al., 2021). ATS research in other languages has received much less attention (Martin et al., 2022). Figure 1 shows that the growth of English text simplification research outpaces progress in other languages.
A diverse multilingual benchmark is essential for a more comprehensive evaluation of multilingual simplification methods, pre-trained models, and evaluation metrics. The lack of a multilingual benchmark that covers a set of high, medium, and low-resource languages belonging to different scripts and language families hinders advancement in multilingual ATS. In this paper, we address this gap in the field by introducing the MultiSim benchmark that covers 27 text simplification datasets (complex-simple pairs) in 12 different languages. MultiSim consists of a collection of datasets from the literature that we unify into a single format for easier accessibility to the research community. In summary, our main contributions are as follows:
We present a comprehensive literature survey of all existing multilingual text simplification corpora, created via several methodologies categorized into four main approaches (§3).
We release the MultiSim benchmark for multilingual text simplification, containing 1,749,056 simple-complex sentence pairs in 12 different languages. To our knowledge, this is the first multilingual benchmark for text simplification. (§4).
We run various experiments using pre-trained multilingual language models and analyze their effectiveness in few-shot learning and cross-lingual transfer for challenging cases of low-resource languages or domain-specific simplification (§5). Our results highlight the benefits of domain and language script match for zero-shot transfer. We find that few-shot prompting large language models produces high-quality simplifications in both high and low-resource languages (§6). We validate these findings with human evaluation (§7).
Related Works
Recently researchers have released several multilingual benchmarks to assess that models work well not only in the high-resource settings where they are trained but in all languages. XTREME-R (Hu et al., 2020) is a multitask benchmark across 50 languages. The benchmark focuses on classification, question answering, structured prediction, and retrieval. Another text classification benchmark is XGLUE (Liang et al., 2020), which covers 11 diverse tasks in 19 languages. Finally, the new XTREME-UP benchmark (Ruder et al., 2023) evaluates 88 under-represented languages on 9 tasks from machine translation to OCR to autocomplete.
Single task multilingual benchmarks exist for NLI (Conneau et al., 2018), QA (Lewis et al., 2020; Longpre et al., 2021), causal reasoning (Ponti et al., 2020), semantic similarity (Vulić et al., 2020), style transfer (Briakou et al., 2021), fact checking (Gupta and Srikumar, 2021), fairness (Chalkidis et al., 2022), stance classification (Zheng et al., 2022), text summarization (Giannakopoulos et al., 2015; Ladhak et al., 2020; Scialom et al., 2020), readability (Naous et al., 2023) and more (Gretter, 2014; Meilicke et al., 2012; Li et al., 2020; Raganato et al., 2020). To date, no such benchmarks exist for multilingual text simplification.
2 Multilingual Text Simplification
The most common approach to automatic text simplification is training a statistical or neural sequence-to-sequence generation model on parallel simplification corpora (§3). Besides in English, researchers have done this in Brazilian Portuguese (Specia, 2010), German (Säuberli et al., 2020; Battisti et al., 2020), Spanish (Štajner et al., 2015; Štajner, 2014), French (Cardon and Grabar, 2020), Japanese (Goto et al., 2015; Maruyama and Yamamoto, 2017), Danish (Klerke and Søgaard, 2013), and Russian (Shatilov and Rey, 2021; Fenogenova, 2021).
Zero-shot and unsupervised learning are promising directions for multilingual text simplification. Mallinson et al. (2020) showed zero-shot simplification in German using a German decoder on a transformer architecture trained on English simplification. Additionally, Martin et al. (2022) showed that mining a massive amount of paraphrases and simplifications as training data was sufficient to achieve state-of-the-art performance in English, Spanish, and French text simplification.
Parallel Simplification Corpora
So far, 31 parallel simplification corpora exist in non-English languages. We organize the discussion of these corpora by their creation strategy. Figure 2 shows the amount of data in each language divided by collection strategy. Table 1 summarizes the details of each corpus stratified by language. We provide more detail on each corpus in Appendix A.
Manual simplification is the most widely-used method for crafting a monolingual parallel simplification corpus. For manual simplification, annotators ranging from experts to crowdsourced workers to the researchers themselves use simplification guidelines (Aluísio et al., 2008b; Siddharthan, 2004; Gonzalez-Dios et al., 2018) to simplify complex documents manually. Researchers have used this methodology to create 17 resources in 10 different languages. Manual simplification gives researchers control over the types of simplification operations in the dataset. However, guiding annotators using rules can result in unnatural simplifications. Stajner (2014) found that relaxing guidelines led to more effective simplifications.
2 Automatic Collection
Manual simplification at scale is a costly and time-consuming process. An alternative approach is to use automatic collection methods, which leverage various online knowledge bases or web sources to increase the size of these resources. However, this can come at the cost of sacrificing quality, as automatically paired sentences may not always be an exact complex-simple match. Additionally, it can be challenging to control the level of simplification.
One common source for automatic collection is Wikipedia, which is available in many languages and often has both original and simplified versions with the same content.https://simple.wikipedia.org/wiki/Main_Page This has been used to build parallel simplification corpora (Jiang et al., 2020; Tonelli et al., 2016; Cardon and Grabar, 2019), where researchers match articles on the same topic from the regular and simple Wikipedia versions. They then leverage automatic aligners such as a neural CRF aligner (Jiang et al., 2020) or CATS (Štajner et al., 2018) to find matching sentences between the two articles.
Other sources of automatically simplified sentences include web scrapes like Common Crawl (Wenzek et al., 2020). For web scrapes, sentence embeddings (Heffernan et al., 2022) are used to find similar sentences to pair up. Then additional filtering is applied to ensure the sentences are not exact matches. A readability measure can be used to ensure that one sentence is simpler than the other. This was the strategy used to create PaCCSS-IT (Brunato et al., 2016) in Italian and MUSS (Martin et al., 2022) in English, Spanish, and French.
3 Machine Translation
Some datasets are machine translations of existing large resources. For example, the French WikiLargeFR (Cardon and Grabar, 2020) and Russian RuWikiLarge (Sakhovskiy et al., 2021) are both machine translations of the English WikiLarge corpus (Zhang and Lapata, 2017). While this allows for large resources in multiple languages, it has two significant drawbacks. Firstly, the final dataset lacks the cultural identity of naturally occurring data in the target language. Secondly, machine translation errors can be introduced in the process, potentially impacting the dataset’s quality.
4 Target Audience Resources
The final monolingual parallel corpora category is resources created for specific target audiences, such as individuals with lower literacy levels or second language learners. These resources typically have the highest quality, but they can be expensive to produce and therefore are relatively rare. There are currently target audience resources available in seven languages.
One company that specializes in creating high-quality target audience resources in English and Spanish is Newselahttps://newsela.com (Xu et al., 2015). Founded in 2013, Newsela is a Series D startup that has attracted over $100 million in funding to support its goal of promoting meaningful classroom learning at all levels. The company employs a team of content producers with extensive teaching experience in K-12 education to train and manage a network of freelance writers who create the simplifications.
Several national news agencies have established systems for creating simplified versions of their articles. For instance, News Web Easyhttps://www3.nhk.or.jp/news/easy/ in Japan is a division of the Japan Broadcasting Corporation, a public media organization. News Web Easy targets Japanese second language learners and primary and secondary school students (Goto et al., 2015). The German government funds the Austrian Press Agency to produce TopEasyhttps://science.apa.at/nachrichten-leicht-verstandlich/ (Säuberli et al., 2020), a simplified version of their news published each weekday. Similarly, the publicly funded Danish Broadcasting Corporation (DR) offers simplified versions of their stories called DR Ligetilhttps://www.dr.dk/ligetil (Straightforward) (Klerke and Søgaard, 2012).
The MultiSim Benchmark
We release the MultiSim Benchmark, a collection of 27 parallel simplification corpora in 12 languages and 4 Scripts. 18 of these corpora are open-sourced and are available online. For 9 corpora, permission must be obtained from the original authors. We provide data loaders for these resources.
We included all languages with open-source parallel sentence-aligned text simplification corpora. This covers eight languages: English (en), French (fr), German (de), Italian (it), Japanese (ja), Russian (ru), Slovene (sl), and Urdu (ur). Six languages have corpora available on request. We provide data loaders and splits to make these resources compatible with the MultiSim benchmark. Such resources exist in Basque (eu), Brazilian Portuguese (pt-br), Danish (da), German (de), Russian (ru), and Spanish (es). Some resources are entirely private due to copyright protection or data-sharing permissions. These resources are in Arabic (ar), German (de), Japanese (ja), and Spanish (es).
2 Domains
The MultiSim benchmark spans 8 domains. Literature sources are simplified versions of novels. Science Communications are popular science articles already written for public consumption and then rewritten at a simpler level. News sources are simplified versions of articles and news stories. Wikipedia sources are pulled from original and simple Wikipedia sites. Website sources generally come from web scrapes like Common Crawl (Wenzek et al., 2020) or specific target websites with original and simplified texts. Medical documents are drug leaflets, clinical notes, and similar texts written for doctors but simplified so that an average patient could understand. Government documents are taken from government policies and simplified to use more common vernacular. Encyclopedic documents are informational texts like Wikipedia but from other encyclopedic sources.
3 Pre-processing and Splitting
For any resource that provided a train, test, dev split, we include the original split of the data in our collection. Otherwise, we randomly divided all sentence pairs into train, test, and dev sets. For resources under 10,000 sentence pairs, we used 80%/10%/10% splits. For resources above 10,000 sentence pairs, we randomly sampled about 1,000 sentences each for the test/dev sets. For resources above 500,000 sentence pairs (WikiAuto), we randomly sampled about 5,000 sentences each for the test/dev sets. We report split sizes in Table 2.
Since several resources in the benchmark come from overlapping domains (i.e., Wikipedia, Web, News), repeat sentences exist between the original datasets. To fix this, we identified overlapping sentences and ensured they fell in the same split by swapping with randomly sampled sentence pairs. We repeated this process until all splits were completely independent.
Experiments
For automatic evaluation, we use SARI (Xu et al., 2016), the average of the F1 score for adding, keeping, and deleting n-grams (). SARI has been shown to correlate with human judgments of simplicity (Xu et al., 2016). We also report BLEU (Papineni et al., 2002), a common metric in machine translation. Although BLEU scores do not measure simplicity (Sulem et al., 2018), we use them as a check for grammatically and meaning preservation (Xu et al., 2016). We compute all evaluation metrics using EASSE evaluation suite (Alva-Manchego et al., 2019).
2 Baselines
To put the results of our experiments in perspective, we compare them with two common baselines.
Identity The original sentence is copied and reported as the simplification. This baseline earns high BLEU scores from the high token overlap between original and simple sentences.
Truncation The last 20% of words are cut from the original sentence. This baseline achieves high SARI scores because it balances keeping/deleting tokens, two operations SARI measures.
3 Models
For fine-tuning we used mT5 Base (Xue et al., 2021) (580M Parameters). We calculated the S-BLEU score between the original and simple sentences in the training set and filtered out all sentences outside of the range as done by Maddela et al. (2021) to remove identical pairs (high BLEU) and misalignments (low BLEU). We also added four control tokens to the input sentence with information about the character-length compression, Levenshtein similarity, word rank, and dependency tree depth of the output following Martin et al. (2020). We used a grid search of control tokens on the dev set to find the combination that yielded the highest SARI for evaluation. We used BLOOM (Scao et al., 2022) (176B Parameters) for a few-shot. We report hyperparameters, prompts, and details for both models in Appendix C.
Results
We evaluated the mT5 models on all 27 datasets and fine-tuned them in 3 different settings. Single: On the training set of the dataset we are testing. Language: On the joint training set of all data in the same language. All: On the joint training set of all data across all languages. We remove results for training sets with fewer than 3,000 sentence pairs after S-BLEU filtering as we found this was not enough training data to unlearn the pre-training objective. We report the results of these experiments in Table 3.
Joint training improves performance in non-English languages. Joint all training improves SARI scores across every language besides English. English already has a wealth of in-language data so it performs best with joint-language training. There are specific datasets where joint-all does not achieve the highest SARI in other languages. Typically these are within one SARI point. Notably, PaCCSS-IT decreases in performance with more data. This may be due to the automatic collection approach to PaCCSS-IT which is prone to collect slightly noisy data. The similar BLEU scores to the identity baseline for all results suggests consistently high fluency.
2 Zero-shot Cross-lingual Transfer
We assess zero-shot cross-lingual and cross-domain transfer by training on one dataset and evaluating on another. We experiment with transfer to a small, domain-specific Italian dataset: Terence, and two low-resource language datasets: CBST in Basque and SimplifyUR in Urdu. The transfer experiment results are shown in Table 4.
Matching script and language improve transfer performance. In transfer experiments to Italian and Basque, we see a notable improvement in BLEU and SARI scores with datasets in matching scripts (Latin, in this case). The best transfer results in Italian come from another dataset in the same language, demonstrating that in-language transfer learning trumps cross-lingual transfer if the data is available. Transferring across scripts typically corresponds to lower performance except when domains match.
Domain match can help regardless of script. In Urdu, we find that the best cross-lingual transfer results come from datasets in the same domain. This is true even though none of the transfer resources are in the Arabic script. In the transfer to Terence (Italian literature corpus), the Russian dataset with the matching domain, RuAdaptLit, outperforms RuWikiLarge, another Russian dataset from a different domain. Still, the domain alone does not guarantee strong transfer performance. EasyJA and NewselaEN performed poorly in Urdu transfer despite matching in domain.
Russian is a good candidate language for cross-lingual transfer. For every test setting, the Russian corpus, RuAdaptLit, transfers well to the target dataset. Additionally, RuWikiLarge transfers better to Terence than the comparable WikiLargeFR, even though both datasets are machine translations of the same English WikiLarge corpus. This suggests that Russian is a good candidate language for cross-lingual transfer, which is in line with the findings of Turc et al. (2021) that Russian is a better choice than English as a pivot language for zero-shot cross-lingual transfer.
3 Prompting Multilingual Language Models
We assess few-shot performance in two settings. Semantic Similarity: We computed LASER sentence embeddings (Schwenk and Douze, 2017) for all sentences in the train, test, and validation set. During evaluation, we used the nearest neighbors in the train set by cosine distance as examples. Random Sampling: We choose random sentences from the train set as examples during evaluation. We highlight interesting findings of our few-shot experiments here. Fewshot results for all datasets are available in Table 8.
Few-shot prompting is promising for low-resource languages. Figure 3 shows semantic similarity sampling few-shot performance for the three low-resource languages in our study: Urdu, Basque, and Slovene. In all low-resource languages tested, few-shot outperforms fine-tuned mT5 trained on all data (+4.62, +5.48, +6.72 SARI, respectively). Within five examples few-shot exceeds fine-tuned performance. With limited resources, few-shot prompting is a good alternative to fine-tuning.
Semantic similarity outperforms random sampling. Figure 4 shows semantic similarity vs. random sampling for few-shot evaluation on four diverse datasets: SimplifyUR (low resource language), EasyJapanese (manually simplified), CLEAR (medical domain), and RuWikiLarge (machine translated). In all cases prompting with semantic search outperformed random examples. This trend persists across languages and domains. Typically more examples improved performance, but any past five had marginal benefits.
Human Evaluation
For manual evaluation, we enlisted eight volunteers to annotate system outputs in English, Russian, Italian, and Urdu (two in each language) for three properties: Adequacy (is the meaning preserved?), Fluency (is the simplification eloquent/grammatical?), and Simplicity (is the output simpler?). This follows the standard annotation methodology in text simplification research (Martin et al., 2022; Xu et al., 2016). We asked annotators to rate 20 sentences for each model using a Likert scale from 1-5. We report the manual evaluation key in Table 9. Since our ratings are ordinal data, we measure annotator agreement using Krippendorff’s alpha Krippendorff (2011) through the Fast Krippendorff Library Castro (2017). We achieve a high-reliability coefficient in all languages suggesting good annotator agreement. Specifically we calculate in English, in Russian, in Italian, and in Urdu.
Table 5 shows the manual evaluation results. Single dataset fine-tuned performance was deficient in Italian and Urdu because these datasets were small resources (1,012 pairs and 736 pairs, respectively). The Joint-all model performed consistently well across all datasets but did not outperform English language training for English. This aligned with our earlier findings (§6.1) and suggests that very large-scale in-language data is better than multilingual training if such data is available. Russian transfer using RuAdaptLit outperformed English transfer to both Italian and Urdu, reinforcing our observation that Russian is a strong choice for cross-lingual transfer (§6.2). We observe the best model scores from five-shot BLOOM to be on par with the reference simplifications in Italian/Russian and scoring slightly below reference simplifications in English/Urdu. This finding suggests that Few-shot prompting is effective for text simplification in both high and low-resource languages. In the low-resource Urdu setting, few-shot prompting yielded the best results, further substantiating our observations from few-shot prompting experiments (§6.3).
Conclusion
We release MultiSim, the first multilingual text simplification benchmark, a collection of 27 sentence-aligned parallel corpora in 12 diverse languages. We collected these resources by surveying the literature for all existing text simplification resources in non-English languages, which were created via distinct methodologies that we categorize into four main approaches. Using MultiSim, we perform fine-tuning, few-shot, and zero-shot cross-lingual transfer experiments with generative multilingual language models (mT5, BLOOM), which revealed new insights in multilingual text simplification. Our results demonstrate the value of domain and script match for zero-shot cross-lingual transfer. We show that Russian is a good candidate pivot language, outperforming transfer from English in two of our case studies on low-resource and out-of-script languages. Further, we show that few-shot prompting BLOOM with examples obtained via semantic similarity outperforms fine-tuned models for low-resource languages. By releasing this benchmark, we hope to encourage and enable the development and evaluation of multilingual models and evaluation metrics for text simplification.
Limitations
This benchmark compiled and analyzed existing resources collected from diverse methods and domains. Although we demonstrated how careful use of these resources could transfer well to other resources, along with a manual analysis of a varied set of corpora, we cannot guarantee the quality of each resource or validate the methods that the original authors used to create them. We explore each dataset’s linguistic properties in Appendix B. However, we encourage a deeper exploration of the quality of individual resources by researchers that speak the 12 languages included in this benchmark and corresponding data loaders.
Additionally, the human evaluation performed in this study was limited in scope and served primarily to validate the findings by automatic metrics. A more extensive evaluation with more annotators evaluating more sentences would be beneficial in order to draw further conclusions.
Furthermore, some of the resources discussed in this paper were automatically aligned. Although Neural CRF models in English have been shown to yield high-quality alignments (Jiang et al., 2020), other alignment algorithms such as TF-IDF scoring (Nelken and Shieber, 2006) have been shown to result in a high number of false positives (Xu et al., 2015). Future work could include realigning automatically aligned corpora using an embedding-based sentence alignment model trained on manually annotated alignment data (Jiang et al., 2020). We will continue updating this benchmark as updates are made to the underlying datasets, and new multilingual resources are released.
Acknowledgments
We thank Yang Chen and Yao Dou as well as three anonymous reviewers for their helpful feedback on this work. We also thank Govind Ramesh, Nour Allah El Senary, Luca Castagna, Lory O’Brien, Franco Paglione, Livia Paglione, Anton Lavrouk, Leah Levin, Irina Levin, Muhammad Hassan Maqsood, and Talha Ahmad Khan for their help with human evaluation. Furthermore, we thank Mounica Madella for sharing her control token scripts. This research is supported in part by the NSF awards IIS-2144493 and IIS-2112633, ODNI and IARPA via the BETTER program (contract 2019-19051600004) and the HIATUS program (contract 2022-22072200004). The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of NSF, ODNI, IARPA, or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copyright annotation therein.
References
Appendix A Resource Summary
Here we provide a brief summary of the 34 monolingual text simplification parallel corpora surveyed in this work. Our summary focuses on the domain, target audience, and collection strategy of each resource.
ASSET (Alva-Manchego et al., 2020) is a high quality collection of 2,390 original sentences from the TurkCorpus (Xu et al., 2016) which sampled from Wikipedia. ASSET contains 10 manually written simplifications for each of the original sentences using a variety of rewrite operations.
The Newsela English corpus (Xu et al., 2015) was produced by professional writers at Newsela, a U.S. company dedicated to providing high-quality simplifications of informational content for schools. Each article was written at 5 levels, including the original article (level 0) and 4 levels of simplification. For this paper we use the neural CRF alignments provided by Jiang et al. (2020).
WikiAuto (Jiang et al., 2020) is a neural CRF aligned corpus of original and simple wikipedia documents. This is currently the largest text simplification resource available with over ten million original sentences. Once aligned, the number of sentence pairs reduces to just under 600,000. This is because most of the original and simple wikipedia articles are not exact rewrites.
A.2 Spanish
The FIRST corpus (Orasan et al., 2013) was collected as a part of the EU-funded FIRST project to develop a tool to assist people with autism spectrum disorder (ASD) in reading and understanding written documents. The corpus contains 25 original and simplified documents from literature, news, health, general culture, and instructions. The simplifications were performed manually by experts with experience working with individuals with ASD.
The Simplext corpus (Saggion et al., 2015) consists of 193 articles from the news outlet, Servimedia, which were manually simplified based on specific simplification recommendations (Anula, 2011). The articles span four domains: national news, international news, society, and culture.
The Newsela Spanish corpus (Xu et al., 2015) was created alongside the Newsela English corpus for the same audience of students. For analysis in this work the corpus was aligned between adjacent levels (i.e., 0-1, 1-2, etc.) using CATS aligner (Štajner et al., 2017).
A.3 Italian
The Terence and Teacher corpora (Brunato et al., 2015) were the first two Italian parallel corpora for ATS. The Terence corpus consists of 32 simplified short stories for children and the Teacher corpus consists of 18 texts from educational websites. The Terence corpus was simplified by experts with target rules, while the Teacher corpus was simplified by teachers targeting second language learners.
SIMPITIKI (Tonelli et al., 2016) is comprised of a Wikipedia-based sub corpus and a government-document-based sub corpus. SIMPITIKI Wiki used crowdsourced Wikipedia simplifications by extracting edits with keywords such as ”simplified” from the edit history of an Italian Wikipedia dump. Prior work has shown that the quality of Wikipedia simplifications is not guaranteed (Xu et al., 2015), however, the simplifications were manually selected to ensure quality. This is the primary corpus studied as Simpitiki throughout this paper. SIMPITIKI PA was produced by the authors for comparing simplification operations and was later absorbed as a subset of the AdminIT corpus (Miliani et al., 2022). The authors manually simplified public administration documents about building permits and kindergarten admittance.
PaCCSS-IT (Brunato et al., 2016) instead extracted simplification pairs from a large corpus of text. The researchers assumed that in a large enough text dataset (i.e., scraped from the internet), both complex and simple sentences that have similar meanings were bound to exist. Brunato et al. (2016) found over 63,000 such matches using cosine similarity and an SVM using lexical, morpho-syntactic, and syntactic features. Upon manual analysis by the authors, it turned out that about 85% of the pairs were aligned correctly, while 74% of those correct pairs were actually simplifications.
A.4 French
CLEAR (Cardon and Grabar, 2019) is a parallel corpus of biomedical texts written in French and automatically aligned using a random forests classifier of 10 textual features. In a manual assessment of 30 documents, the authors found that 98.75% of alignments were correct suggesting a high-precision sentence alignment.
WikiLargeFR (Cardon and Grabar, 2020) is a machine-translated version of the English WikiLarge corpus (Zhang and Lapata, 2017) using OpenNMT-py (Klein et al., 2017). It has 297,753 sentence pairs but different exact counts of complex and simple sentences when accounting for sentence splitting. It was created as a comparison with CLEAR for biomedical text simplification.
Alector (Gala et al., 2020) contains expert simplified versions of 79 texts selected at a 2-4th grade reading level. The authors showed that simplifying the texts reduced misreadings in dyslexic and low literacy readers.
A.5 Japanese
Japanese News (Goto et al., 2015) is created from Japan Broadcasting Corporation (NHK)’s online service NEWS WEB EASY which provides original news articles rewritten by Japanese instructors for simplicity. Goto et al. (2015) used a dynamic programming aligner to align 10,651 sentences. They also manually aligned 2,735 sentences.
EasyJapanese (Maruyama and Yamamoto, 2018) was created for the purpose of improving Japanese resources for foreign citizens. Since Japanese language learners usually have a limited vocabulary, the authors decided to produce a parallel corpus using a vocabulary of 2,000 common Japanese words. 5 students in the lab manually simplified 50,000 sentences with S-BLEU (Papineni et al., 2002) scores between simplifications of the same sentence ranging from 0.58 to 0.63. The corpus was built from a previous bilingual web crawl of Japanese and English news articles called the Tanaka corpus (Tanaka, 2001).
EasyJapaneseExtended (Katsuta and Yamamoto, 2018) included 34,400 more sentences from the Tanaka corpus (Tanaka, 2001) with simplifications crowdsourced from the CrowdWorkshttps://crowdworks.jp/ platform. The authors measured the S-BLEU scores on 100 sentences that each of the 7 workers simplified and found that for 70% of the workers, S-BLEU scores exceeded 0.4.
A.6 Brazilian Portuguese
PorSimples (Aluísio and Gasperin, 2010) was one of the first ATS projects. It had the express purpose of simplifying texts for individuals with reading difficulties. For this project, Caseli et al. (2009) collected a parallel corpus of 104 news articles. The articles were simplified at 2 levels by a linguist specializing in text simplification. In “natural” simplifications, the linguist could choose how to simplify the text. In “strong” simplifications, the linguist had to follow very specific rules (Specia et al., 2008; Aluísio et al., 2008a).
A.7 German
Simple German (Battisti et al., 2020) started from Klaper et al. (2013)’s work that scraped texts from the internet and aligned them using a monolingual sentence alignment algorithm of Barzilay and Elhadad (2003). Battisti et al. (2020) further improved upon this with much more data and better alignment algorithms of CATS (Štajner et al., 2018) and MASSAlign (Paetzold et al., 2017). The original paper reports 378 documents with 17,121 original and 21,072 simple sentences. Note that, these numbers differ from those in Table 1 as the availability of online articles has changed since the original publication.
TextComplexityDE (Naderi et al., 2019) was created to measure text complexity in German. 1000 sentences were taken from German Wikipedia and 100 sentences from Simple German (Klaper et al., 2013). German second language learners rated the sentences on a 7-point Likert scale for complexity. The 250 most complex sentences were manually simplified by native speakers.
GEOLinoTest (Mallinson et al., 2020) was built as an evaluation dataset for a zero-shot ATS model. Mallinson et al. (2020) extracted 20 articles about nature, physics, and people from GeoLino, a children’s magazine. A German linguist simplified them to a five to seven-year-old reading level.
German News (Säuberli et al., 2020) contains 3,616 sentences simplified by the Austrian Press Agency on politics, economy, culture, and sports. The authors trained neural simplification models on the corpus and found success, including a simple sentence matched to itself during training (Palmero Aprosio et al., 2019).
The Klexikon corpus (Aumiller and Gertz, 2022) is a mapping of documents from German Wikipedia to the children’s encyclopedia site: Klexikon https://klexikon.zum.de. Klexikon targets German readers aged six to twelve. Like WikiAutoEN, Klexikon is a large-scale alignment of documents, however, unlike similar resources, Klexikon does not yet have a gold-standard automatic sentence alignment. The authors of Klexikon are working to release this alignment which will make Klexikon a very large German text simplification resource.
Simple Patho (Trienes et al., 2023) is an upcoming biomedical text simplification corpus for German of 851 clinical reports simplified by nine medical students. Due to privacy concerns, the dataset is not yet available. When it is released, it will serve as a large and high-quality medical text simplification corpus for the community.
A.8 Basque
The Corpus of Basque Simplified Texts (CBST) (Gonzalez-Dios et al., 2018) contains 227 sentences from 3 science popularization documents simplified to two distinct levels. Two people simplified the documents: a translator without simplification experience who focused on simplification guidelines (Mitkov and Štajner, 2014) and a language teacher who focused on intuitive transformations.
A.9 Danish
DSim (Klerke and Søgaard, 2012) extracted 3,701 pairs of news telegrams from the Danish Broadcasting Corporation’s news and educational services. The corpus was simplified by journalists to help reading-impaired adults and adult learners of Danish. The documents were automatically sentence-aligned using TF-IDF scores.
A.10 Urdu
SimplifyUREval (Qasmi et al., 2020) was a corpus made for evaluating the ATS model SimplifyUR. The model used word substitutions to propose simplifications to Urdu. The evaluation corpus contains 500 sentences from newspapers, magazines, books, and literary journals manually simplified by a linguist with a doctorate in Urdu. Two additional native Urdu speakers manually verified 50 sentences and had an inter-annotator agreement of 0.9, measured by Cohen’s Kappa.
A.11 Russian
RuWikiLarge (Sakhovskiy et al., 2021) is a machine translated version of EnWikiLarge (Zhang and Lapata, 2017). It has 248,111 sentence pairs but different exact counts of complex and simple sentences when accounting for sentence splitting. It was created as a resource for the RuSimpleSentEval shared task (Sakhovskiy et al., 2021).
RuSimpleSentEval (RSSE) (Sakhovskiy et al., 2021) was mined from Russian Wikipedia’s most popular pages also for the RuSimpleSentEval. Crowdsource workers on Yandex Toloka were asked to simplify the sentences.
RuAdapt (Dmitrieva and Tiedemann, 2021) was created from 6 collections of several novels each and 16 individual classical and modern Russian literature books. The simplified books were prepared by Russian-as-a-Foreign-Language (RaaFL) teachers. The corpus was aligned with Bleualign (Sennrich and Volk, 2010) and CATS (Štajner et al., 2017). In addition, researchers contributed both encyclopedic simplifications and fairytale simplifications (Dmitrieva et al., 2021).
A.12 Slovene
The SloTS corpus (Gorenc and Robnik-Šikonja, 2022) pulls from 10 existing texts simplified by the RISA Institute. The RISA Institute is an organization that publishes easy-to-read Slovenian novels and news. SloTS is a collection of about 1,000 sentence pairs sampled from 10 novels and manually aligned between the original and simplified versions. This dataset was used to train a Slovene text simplification model built on SloT5 (Ulčar and Robnik-Šikonja, 2022).
A.13 Arabic
Saaq al-Bambuu (Al-Sanousi, 2016) is an internationally acclaimed Arabic novel that has been rewritten for Arabic-as-a-second-language learners. Khallaf and Sharoff (2022) sampled 2,980 parallel sentences from the original and simplified books at two different levels. Unfortunately, due to copyright restrictions, the corpus is not available publicly.
Appendix B Corpus Analysis
This section provides some key statistics of the corpora introduced in Section 3.4. In performing this analysis, we hope to highlight the differences between various corpora and offer greater insight into their quality and composition to facilitate future research.
The Basic Statistics computed were vocab size, token count, average tokens per sentence, average characters per token, and average sentences per doc. All of the statistics besides average sentences per document depend heavily on the word tokenization of the corpus. For space-delimited languages, we used the Toktok tokenizer (Dehdari, 2014) from the natural language toolkit (NLTK) (Bird et al., 2009). For our work, this included all languages besides Urdu and Japanese. For Urdu tokenization, we used UrduHackhttps://docs.urduhack.com/en/stable/, a Python library built for academic researchers and professional developers working on Urdu NLP projects. For Japanese, we used fugashi (McCann, 2020), a tool for Japanese tokenization in Python. We used the unidic dictionary (Den et al., 2008) to define the Japanese vocabulary for tokenization.
Basic statistics results are reported in Table 6. In general, the trend between original and simplified texts was reduced vocab size, reduced token count, reduced sentence length, and reduced word length. This corresponds well with the expected simplification operations of replacing longer, more complicated words with shorter, more common words. It also aligns with a common simplification strategy of splitting longer sentences into two or more short sentences. This explains why the average sentence length decreased but in many document-aligned corpora, the average sentences per document increased. There were some exceptions to this. Notably, the Japanese corpora had a higher token count and higher average tokens/sentence. This could be due to the limited vocabulary used when creating these two corpora. As a part of the Easy Japanese corpus creation authors were limited to just 2,000 Japanese words. The fugashi tokenizer identified more than 2,000, but still reported a large drop in vocab size. The authors may have needed to be creative with how they chose to rewrite sentences using the limited vocabulary. This could’ve led to many edits where the author explained a complex idea in several simple words instead of using one more complicated word. Another easily explained set of outliers is SimplifyUR and RSSE having larger vocab sizes from simple to complex. Both of these corpora allow multiple translations of the same original sentence. This means for a given sentence pair with the same original sentence the original vocab size will remain the same while the simple vocab size might increase.
B.2 Document-level Compression
In order to measure the editing levels on a document scale, we investigated document-level compression. Document-level compression is the ratio of the number of characters in the simple document to the number of characters in the original document (Xu et al., 2015). A low document compression ratio indicates a lot of deletions between the original and simplified text, while a high document compression ratio suggests lengthier operations like sentence splitting or rephrasing.
We report the compression ratios of all document-aligned corpora in Figure 6. Most of the compression ratios are approximately normally distributed. For many of the corpora, the compression ratio is centered around one, meaning the original and simplified documents are about the same length. This closely matches the low edit ratios that many of the corpora have (see section B.4). A few of the corpora (Simplext, Newsela, and German News) have lower means instead suggesting more significant document-level edits, such as deletion of entire sentences.
B.3 Sentence-level Edit Operations
Edit operations describe the types of simplifications that were performed to transform from an original sentence to a simplified sentence. These can only be computed for corpora that are sentence aligned. There are 6 edit operations we tracked from the alignments. Each operation corresponds to a mapping (:) of original sentences to simple sentences. The operations are deletion (1:0), split (1:n), same (1:1), change (1:1), merge (n:1), and insert (0:1). To determine the difference between “same” and “change”, the Levenshtein distance (Levenshtein, 1965) was measured between the original and simplified sentence. This distance was divided by the length of the longer sentence. If the difference was greater than 5% then the sentences were marked as changed, otherwise, they were considered the same. Levenshtein distance was calculated using the fuzzywuzzyhttps://github.com/seatgeek/thefuzz library.
Table 7 shows the distribution of edit operations. Sentence-level edit operations were reported for both document-aligned corpora as well as sentence-aligned corpora that used sentence splitting (1:n mapping). The most common edit operation across corpora was changing the original sentence, followed by keeping the same sentence, then splitting, deleting, and merging. Interestingly, Spanish corpora, like English ones Xu et al. (2015); Jiang et al. (2020), had more deletion operations than most of the other languages. For the sentence-only aligned corpora, this was because original sentences without simplifications were not included, but this was true even amongst document-aligned corpora.
B.4 Character-level Edit Distances
We analyzed the edit distance distribution (Vásquez-Rodríguez et al., 2021) on all corpora with sentence-level alignments to understand the strength of the edits at a sentence scale. Low edit distances indicate smaller simplifications while high edit distances indicate big changes. We again used character-based Levenshtein distance to measure edit distance. The Levenshtein distance was divided by the length of the longer sentence to obtain a ratio from zero to one. Zero meant the sentence wasn’t edited at all, while one meant the sentence was completely different.
The edit distance ratios can be found in Figure 5. For about half of the corpora, the mean edit distance fell below 20%. For the other half mean edit distance ratios ranged from 0.2 to 0.6. Corpora with an approximately normal distribution and higher variance demonstrate a wide variety of both minor and major sentence edits. Corpora with low means and a high concentration of low edit ratios primarily consist of slight modifications.
Appendix C Experimental Details
Fine-tuning For fine-tuning experiments we used the mT5 Base (Xue et al., 2021) architecture (580M Parameters). We used the sentence piece tokenizer (Kudo and Richardson, 2018) and limited inputs/targets to a length of 128 tokens. We used a learning rate of 5e-5 and the AdamW Optimizer Loshchilov and Hutter (2019). Decoding was done using beam search with 4 beams. The train batch size was set to 8. We train for 5 epochs. For any dataset without a development set, we removed 10% of the training set up to 1,000 sentences to create our own dev set. The training was performed on three NVIDIA A40 GPUs.
We also perform preprocessing on the training data inputs. We compute sentence BLEU scores between the original and reference simplifications. For any sentence pairs with an S-BLEU score below 10 or above 70, we remove it from the training set (Maddela et al., 2021). This helps reduce both misaligned and identical pairs. Following Martin et al. (2020) we also add control tokens to the input sentences. We include a character length compression token
Fewshot For fewshot experiments we used BLOOM (Scao et al., 2022) (176B Parameters). We used the HuggingFace inference APIhttps://huggingface.co/inference-api to prompt BLOOM with sampling. We used a temperature of 1.0 and a repetition penalty of 0.0. We formatted prompts to BLOOM as:
The prefixes ”Original” and ”Simple” were always in English. Outputs were appended to the prompt and repeated back to the model until the output contained an end quotation followed by a new ”Original:”. This was to prevent half-completed simplifications.
Manual Evaluation Table 9 shows the key provided to the human annotators in each language. Annotators were volunteers with fluency in the target language for annotation. We randomly sampled 20 sentences from the test set of each of the four datasets and we used these 20 sentences to compare across all models. For the ”Reference” baseline if any sentence had more than one possible reference simplification we randomly sampled a reference. For any of the system outputs, if the sentence was completely nonsense annotators were instructed to rate the sentence with a score of 1 on all aspects.
All annotators were volunteers that were informed that we were ”measuring the quality of various machine learning models that have been trained/prompted to simplify text”. All of the volunteers were told they were assessing sentences to be used in evaluation for a research project. For any student employees that volunteered their time for evaluation, they were paid at their normal hourly rate of $18 per hour. Other colleagues that volunteered their time did so on a strictly voluntary basis.