ReadMe++: Benchmarking Multilingual Language Models for Multi-Domain Readability Assessment
Tarek Naous, Michael J. Ryan, Anton Lavrouk, Mohit Chandra, Wei Xu
Introduction
Automatic readability assessment is the task of determining the cognitive load needed by an individual to understand a piece of text Vajjala (2021). Assessing the readability of a sentence is helpful for many applications, including controlling the complexity of machine-generated text Chi et al. (2023); Agrawal and Carpuat (2019), ranking search engine results by their readability level Fourney et al. (2018), or developing tools such as Grammarly that assist writers in enhancing the quality of their text. Enabling such technologies for all the world’s languages requires readability prediction methods that generalize across language families and text genres.
Recent works in multilingual transformer language models such as mT5 (Xue et al., 2021) and XLM-R (Conneau et al., 2020) enable the development of language-agnostic multilingual readability measurements (Martinc et al., 2021). However, without a high-quality, diverse multilingual benchmark there is no way to accurately assess and compare the supervised, unsupervised, and few-shot methods that use these newly developed large language models. Current evaluation resources for sentence readability assessment suffer from a few crucial shortcomings. First, existing datasets primarily contain sentences from Wikipedia Naderi et al. (2019); Arase et al. (2022); Štajner et al. (2017) and news articles Brunato et al. (2018). However, it is critical for readability assessment methods to perform well in any textual domain. Language models have been shown to struggle when handling data from a different domain outside of their training corpus Plank (2016); Farahani et al. (2021); Arora et al. (2021). Hence, a domain-diverse benchmark is essential in assessing model domain generalization. Additionally, researchers often utilize document-level readability data as an approximation for sentence-level readability (§2), due to a lack of human readability ratings on individual sentences Martinc et al. (2021); Lee and Vajjala (2022).
To address these gaps in research, we release ReadMe++, a diverse multi-domain dataset for multilingual readability assessment (Figure 1). Our sentence-level human-annotated dataset draws sentences from 112 distinct sources in 5 languages, enabling a more holistic evaluation of model domain generalization and cross-lingual capabilities in readability assessment. In summary, our contributions are as follows:
We release ReadMe++, a sentence-level human-annotated dataset of 9757 sentences from 112 data sources on varied topics for five languages: Arabic, English, French, Hindi, and Russian, covering four language scripts (Arabic, Latin, Brahmic, Cyrillic) (§3.1). The language diversity in ReadMe++ encourages the study of under-explored challenges like the effect of transliterations on readability in non-Latin scripts (§4.5)
We benchmark both monolingual and multilingual language models for multi-domain readability assessment and show that fine-tuned models outperform few-shot prompting and unsupervised approaches (§4.4), highlighting the usefulness of human annotations for readability.
We show that models trained using ReadMe++ perform better across domains and exhibit superior cross-lingual transfer capabilities from English to six target languages: Arabic, French, Hindi, Russian, Italian, and German compared with models trained on previous datasets (§4.5).
Related Work
Many datasets used in readability assessment research have only document-level ratings, as they were directly collected from data sources (e.g., textbooks) that provide parallel or non-parallel text at varied levels of writing. These include WeeBit Vajjala and Meurers (2012), Newsela Xu et al. (2015), Cambridge Xia et al. (2016a), OneStopEnglish Vajjala and Lučić (2018), VikiWiki Azpiazu and Pera (2019), Slovenian SB Martinc et al. (2021), English-Chinese LR Rao et al. (2021), ALC Khallaf and Sharoff (2021), Gloss Khallaf and Sharoff (2021), and Philippine Corpus Imperial and Kochmar (2023). While appropriate for assessing document-level readability, such datasets are less accurate for sentence-level readability assessment than the datasets that have ground-truth readability labels for each individual sentence Arase et al. (2022); Cripwell et al. (2023).
Sentence-based Readability.
Only a few existing datasets De Clercq and Hoste (2016); Štajner et al. (2017); Brunato et al. (2018); Naderi et al. (2019) were created by manually annotating individual sentences for their level of readability (see Table 1). However, these sentence-level annotated datasets are largely limited to high-resource English and European languages that use the Latin script. They are also limited to one or a few data sources while being annotated based on varied scales and, thus are insufficient for studying the robustness of readability measurement methods across text domains. One of the most recent and related works to ours is CEFR-SP Arase et al. (2022), but it only contains English sentences from Wikipedia, Newsela (Xu et al., 2015, leveled news articles), and SCoRE (Chujo et al., 2015, textbooks for learning English). In comparison, our work highlights the importance of both language and domain coverage, resulting in more data diversity (see Figure 2). Our ReadMe++ corpus covers 112 different data sources and is manually annotated at the sentence level in 5 languages. We also adopt the CEFR standard Council of Europe (2001) used by Arase et al. (2022) and the Rank-and-Rate annotation approach used by Maddela et al. (2023) to ensure annotation quality.
Multilingual Readability Assessment.
Several efforts have investigated neural approaches for readability assessment in non-English languages Blaneck et al. (2022); Mesgar and Strube (2018); Sun et al. (2020); Chakraborty et al. (2021); Imperial et al. (2022); Le et al. (2018); Imperial and Kochmar (2023); Azpiazu and Pera (2019); Martinc et al. (2021). Some works have explored cross-lingual transfer from English to French/Spanish Lee and Vajjala (2022) and Chinese Rao et al. (2021). The majority of prior works have been evaluated on datasets with only document-level annotations and/or datasets that cover only a few text domains. Using our dataset, we benchmark the capabilities of recent large language models to generalize to a large variety of text domains in diverse language scripts (i.e., Arabic, Latin, Brahmic, and Cyrillic). We show that models trained using the English portion of ReadMe++ perform better cross-lingual transfer to 6 target languages compared with models trained on previous datasets.
Constructing ReadMe++ Corpus
We present the detailed procedure for constructing the ReadMe++ corpus. To maximize the diversity of topics and genres (Table 2), we identified 112 data sources that are either with open licenses or shareable for non-commercial purposes. A total of 9757 sentences (1945 Arabic, 1669 French, 2861 English, 1524 Hindi, 1758 Russian) were sampled from these sources and then double-annotated. ReadMe++ can support multilingual, cross-lingual, and cross-domain experiments (§4).
The data collection process varies per source but can be categorized into four approaches: (1) obtaining content directly from a website (e.g., Wikipedia), (2) extracting text from sources in PDF format (e.g., contract templates, reports, etc.), (3) sampling text from existing data sources (e.g., dialogue, user reviews, etc.), or (4) manually collecting sentences (e.g., dictionary examples, etc.). Full collection details for each domain and language are provided in Appendix A. For each domain, we collected the available texts from one or more data sources and then sampled 50 paragraphs per domain. For domains collected from highly unstructured sources, such as PDF files, the sampling rate was increased to 100 since it is highly likely that samples may contain text that is not useful for annotation (e.g., headers, titles, references, etc.). Finally, from each paragraph, we sample one sentence that will be used for readability annotation. For quality control, we perform manual post-sampling quality check to filter out any low-quality sentences and sentences that contain toxic or offensive language.
Considering the Influence of Contexts.
In addition to the sampled sentences, we collect up to three preceding sentences as context if available. Many of the sampled sentences could be placed in the body of a paragraph. Some may require context to be fully understood. By providing optional context, we ensure annotators will not mark a sentence as confusing and not easily readable simply because they do not know the context in which it appears. Such cases have not been adequately considered in previous work; for example, Arase et al. (2022) also noticed this problem but avoided it by collecting only the first sentence in a paragraph.
2 Readability Annotation
Previous works on sentence-level readability have used various rating scales such as 0-100 De Clercq and Hoste (2016), 3-point Štajner et al. (2017), or 7-point Naderi et al. (2019); Brunato et al. (2018) scales. However, these scales lack clear readability grounding which causes high levels of annotator subjectivity. Instead, following Arase et al. (2022), we adopt an international language standard, namely the Common European Framework of Reference for Languages (CEFR), which defines the language ability of a person on a 6-point scale (1(A1), 2(A2), 3(B1), 4(B2), 5(C1), 6(C2)), where A is for basic, B for independent, and C for proficient. Each level of the scale is defined by detailed descriptions of what form of text the person can understand (Appendix B). This makes the CEFR scale a good guideline for readability annotation, and help the annotators be more consistent as shown by Arase et al. (2022), compared to prior works that used the Likert scale without detailed definitions.
Rank-and-Rate Annotation.
Rating each sentence independently on a scale of readability comes with the drawback of annotators eventually not differentiating between different sentences. This results in most samples being labeled within one or two levels, limiting their usefulness for statistical analyses McCarty and Shrum (2000). Instead of rating alone as in prior works, we utilize a rank-and-rate approach for readability annotation which mitigates independent sentence rating issues by providing comparative texts. We randomly group sentences into batches of 5 and ask annotators to first rank sentences of a batch from most to least readable and then rate each sentence on the 6-point CEFR scale. By comparing and contrasting sentences within a batch, annotators can better differentiate between the readability of different sentences and produce less subjective ratings. A pilot study was conducted where annotators labeled using rating-alone and the rank-and-rate framework. Overall, annotators expressed a better experience with the rank-and-rate framework and achieved slightly higher agreements than rating-alone in the pilot study. Our interface is shown in Appendix F.
Inter-annotator Agreement.
We recruited two speakers in each language for annotation, all with college-level education. Before annotation, training sessions were conducted to familiarize the annotators with the CEFR levels and the annotation framework. We measure Krippendorff’s alpha () between annotators, obtaining good agreement levels that range between to .
Benchmarking Automatic Sentence Readability Measurements
As shown in Figures 3 and 2, the new ReadMe++ corpus offers a diverse mix of many topics and readability levels, making it an ideal testbed for evaluating automatic readability assessment. We benchmark supervised, unsupervised, and few-shot approaches using recently developed language models. We find that supervised models can reach over 0.8 Pearson correlation with human ratings, but unsupervised and few-shot approaches lag behind.
In the supervised setting, we fine-tune language models to classify sentence readability.
We use mBERT Devlin et al. (2019) and XLM-RoBERTa Conneau et al. (2020) multilingual models and compare to monolingual models that include BERT Devlin et al. (2019) for English, AraBERT Antoun et al. and ArBERT Abdul-Mageed et al. (2021) models for Arabic, CamemBERT for French Martin et al. (2020), and RuBERT Kuratov and Arkhipov (2019) for Russian. For Hindi, we fine-tune MuRIL Khanuja et al. (2021), a model pre-trained on 12 different Indian languages. We also compare to seq-to-seq transformers including mT5 Xue et al. (2021) and AraT5 Elmadany et al. (2022). Model details are summarized in Appendix D.1. We fine-tune all models for 20 epochs using the cross-entropy loss and the Adam optimizer and tune the learning rate in the set . We use the same random train/valid/test split (detailed statistics in Appendix D.2) based on the 80/10/10% ratio per domain for all experiments, except the domain generalization study in §4.5. We select checkpoints based on the best performance on the validation set. We report the average of the results for fine-tuning each model using 5 different random seeds for initialization.
2 Unsupervised Methods
In the unsupervised setting, we leverage the language model distribution to compute a readability score without training. We also compare with several traditional length-based readability formulas.
The Ranked Sentence Readability Score (RSRS) proposed by Martinc et al. (2021) combines neural language model statistics with the average sentence length as lexical feature. It computes a weighted sum of the individual word losses as follows:
where is the sentence length, is the rank of the word after sorting each Word’s Negative Log Loss (WNLL) in ascending order. Words with higher losses are assigned higher weights, increasing the total score and reflecting less readability. is equal to 2 when a word is an Out-Of-Vocabulary (OOV) token and 1 otherwise, assuming that OOV tokens represent rare, difficult words and thus are assigned higher weights by eliminating the square root. The WNLL is computed as follows:
where is the distribution predicted by the language model, and is the empirical distribution where the word appearing in the sequence holds a value of 1 while all other words have a value of 0.
Traditional Readability Metrics.
We use a few common traditional readability metrics based on word and sentence lengths. Specifically, we use the Average Sentence Length (ASL), Automated Readability Index (ARI) Smith and Senter (1967), Flesch-Kincaid Grade Level (FKGL) Kincaid and Robert Jr (1975), and Open Source Metric for Measuring Arabic Narratives (OSMAN) (El-Haj and Rayson, 2016). ARI and FKGL were originally designed for English. Detailed formulas for all these metrics are listed in Appendix C.
3 Prompting Methods
We also evaluate in-context learning using Llama2-7B Touvron et al. (2023) and GPT-4. We provide the model with a definition of readability and the same description of the six CEFR levels (Appendix B) provided to our human annotators. We show the model five randomly sampled in-context examples from the train set and their corresponding CEFR levels, then ask the model to assess the readability of a new sentence based on the CEFR scale provided. Prompt details can be found in Appendix D.3.
4 Main Results
Figure 4 shows the results of the fine-tuned and prompted models. The results achieved by unsupervised methods are shown in Figure 5. We report the Pearson Correlation () between the predictions and the ground-truth labels (additional metrics such as F1 scores are provided in Appendix E.1).
The performance of Llama-2 and GPT-4 with 5-shot examples was lower than that of fine-tuned models in all languages. Fine-tuned models were able to achieve high correlation levels in the range of 0.7 to 0.9, with larger-sized models showing improved performance in nearly all cases, except for XLM-RL which we found to be more sensitive to random initialization. Overall, mT5L was among the best-performing fine-tuned models across all languages. We further analyze how in-context examples impact few-shot performance in § 4.5.
Unsupervised LM-based prediction outperforms traditional length-based metrics.
as shown in Figure 5, LM-based RSRS scores achieve better correlation than traditional readability metrics in all languages, highlighting the usefulness of large language models for readability assessment. We find that RSRS with monolingual models achieves a noticeably lower correlation for non-English languages compared to multilingual models, specifically in languages with non-Latin script (Arabic, Hindi, Russian). We further investigate this gap in § 4.5. Little to no difference is observed in performance in the unsupervised setting between base and large versions of the models.
Fine-tuned models outperform LM-based unsupervised prediction.
In Figure 6, we directly compare the unsupervised and supervised methods. We compute the macro F1 for unsupervised metrics by searching for the optimal thresholds for each metric that maximize the F1 score on the validation set. There is a notable gap in performance between unsupervised and supervised methods, with fine-tuned models outperforming unsupervised metrics. While promising, better unsupervised methods are needed in future work to bridge this gap.
5 Further Analyses
To test how well models perform on unseen domains during fine-tuning, we create new train/val/test splits from ReadMe++ by randomly removing an increasing number randomly sampled domains from the dataset (Table 3). We use the sentences from the removed domains as the test set and use the rest of the dataset for training and validation. For direct comparison, we randomly sample the same amount of train/val sentences in each experiment from the open-sourced Wikipedia-based portion of CEFR-SP Arase et al. (2022) to fine-tune mBERT models. We evaluate these models on the unseen domains test set from ReadMe++. Results are shown in Table 3. Models fine-tuned using ReadMe++ achieve good performance on unseen domains and significantly outperform models trained using CEFR-SP, demonstrating the advantage of data diversity in ReadMe++.
We perform the same experiments in Arabic by comparing to the ALC Corpus Khallaf and Sharoff (2021), which is labeled on 5-scale CEFR levels (A1, A2, B1, B2, C). We convert the labels in ReadMe++ to the same scale of ALC Corpus by combining C1 and C2 into C and then perform a 5-way classification. We observe the same trend, where models trained using the Arabic portion of ReadMe++ achieve good performance on unseen domains and outperform models trained on ALC.
Models trained on ReadMe++ perform better cross-lingual transfer.
We perform zero-shot cross-lingual transfer from English to Arabic, Hindi, Italian, and German by fine-tuning multilingual models using the English subset of ReadMe++. For comparison, we also fine-tune on the same number of train/valid sentences that we randomly sample from the open-sourced Wikipedia-based portion of CEFR-SP Arase et al. (2022) and the full English CompDS Brunato et al. (2018) corpora. We evaluate on the Arabic and Hindi test sets from ReadMe++, as well as Italian CompDS Brunato et al. (2018) and German TextComplexityDE Naderi et al. (2019). Since CompDS and TextComplexityDE rate on scales from 1-7 instead of 1-6 but have only a few level-7 sentences, we merged their level 6 and 7 together. Results are shown Table 4 for XLMRL. The model fine-tuned using ReadMe++ performs much better cross-lingual transfer than models fine-tuned using CEFR-SP or CompDS across all tested languages, reaching high correlation values of 0.7 in most. In several cases, training on ReadMe++ leads to a 50% increase in F1 score and double the correlation value. This trend is also observed across several models which we report in Appendix E.3.
Domain diversity of in-context examples improves few-shot performance.
We study the effect of having domain diversity of in-context examples on the performance of few-shot prompting. We prompt Llama2 by sampling examples for 1, 2, 4, and 8 domains. The domains from which in-context examples are sampled are also randomly sampled for each test sentence. The average correlation from 5 runs is shown in Figure 7, for an increasing number of shots. The impact of increasing domain diversity is clearly shown, with correlation increasing in nearly all cases.
Unsupervised LM-based methods struggle with transliterations.
The RSRS method (§4.2 Eq. 1) assumes that all unseen words by the model’s tokenizer are rare, difficult words that should be assigned higher weights. However, these could also be transliterations of names of new figures in politics, emerging diseases, or even historical names the model never saw during pre-training. We hypothesize that this design choice in RSRS degrades its performance on non-Latin languages since many of these transliterated words do not add to the difficulty level of the sentence for humans. To test this hypothesis, we asked Arabic, Hindi, and Russian annotators to indicate if a sentence contains transliterated words when performing rank-and-rate annotation. This resulted in 320 sentences with transliterations in Arabic (16.45% of Arabic data), 561 sentences in Hindi (36.81% of Hindi data), and 120 sentences in Russian (6.82% of Russian data). We penalize the RSRS scores of those sentences by the factor , where is a penalty factor and is the length of the sentence. We compute the resulting correlation with human annotations for an increasing penalty () to analyze whether decreasing those scores results in higher correlation, since we assume transliterations cause RSRS scores to be unreasonably high. The results are show in Figure 8 for 0.1 increments of . The trends clearly corroborate with our hypothesis; the correlation increasing as the penalty becomes higher up to a certain level. The improvement reaches up to 6-7% for monolingual and Indian models, and up to 1-3% for multilingual models, suggesting that multilingual models are more robust to transliterations, yet it degrades performance for monolingual models. These observations suggest careful consideration for transliterations be given in future research.
Conclusion
We presented ReadMe++, a multi-domain dataset for multilingual sentence readability assessment. ReadMe++ provides 9757 sentences in Arabic, English, French, Hindi, and Russian that are collected from 112 different data sources and annotated by humans on a sentence-level according to the CEFR scale. We showed that models trained using ReadMe++ achieved strong sentence readability assessment performance across different textual domains and performed well in zero-shot cross-lingual transfer.
Limitations
Readability assessment is a general task that can be further customized for a target audience such as children (Lennon and Burdick, 2004) or adults with intellectual disabilities (Feng et al., 2009). In this work, we adopt the CEFR standards which target second-language learners and have been used in many prior works (Arase et al., 2022; Xia et al., 2016b, inter alia) on readability measurements.
Ethical Considerations
We are committed to upholding ethical standards in constructing and disseminating the ReadMe++ corpus. To ensure the integrity of our data collection process, we have made our best effort to obtain data from sources that are available in the public domain, released under Creative Commons or similar licenses, or can be used freely for personal and non-commercial purposes according to the resource’s Terms and Conditions of Use. These sources include user-generated content on public domain books, publicly available documents/reports, and publicly available datasets. We use a small number of randomly sampled sentences for academic research purposes, specifically for labeling sentence readability. We have included a full list of licenses and terms of use for each source in Appendix G. We would like to note that a couple of sources we used require access permission from the original authors, including i2b2/VA Uzuner et al. (2011) and Hindi Product Reviews Akhtar et al. (2016). Therefore, sentences and annotations from these sources will not be shared with the community unless access permission has been obtained from the original authors.
When collecting sentences from social media and forums, we have excluded any sampled sentences containing offensive/hateful speech, stereotypes, or private user information. All annotators were student employees paid at the standard student employee rate of $18 per hour for their time. Every annotator was informed that their annotations were being used to create a dataset for readability assessment.
Acknowledgments
The authors would like to thank Nour Allah El Senary, Govind Ramesh, and Ryan Punamiya for their help in data annotation. This research is supported in part by the NSF awards IIS-2144493 and IIS-2112633, ODNI and IARPA via the HIATUS program (contract 2022-22072200004). The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of NSF, ODNI, IARPA, or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copyright annotation therein.
References
Appendix A More details about ReadMe++
This section provides a description of how sentences were collected from each of the 64 domains of ReadMe++. Table 13 shows statistics of the corpus and Table 14 summarizes the sources from which data was collect for each domain in each language, including publicly available web resources or open-source datasets.
Wikipedia: Wikipedia is an attractive source of multilingual text since most articles are available in a large number of languages. Further, articles belong to a variety of topics where writing style and technicality differ significantly. We select 9 Wikipedia topics and, from each, randomly sample 5 different articles that discuss a certain sub-topic within that topic. For example, an article on “Information Theory” belongs to the “Technology” topic. We scrape the Arabic, English, French Hindi, and Russian versions of each article.
News Articles: We leverage resources used for news category classification research, which we find publicly available datasets for in Arabic Alfonse and Gawich (2022) and English Misra (2022). No similar public resource was found for the other languages.
Research: We collect text from medical, law, politics, and economics research papers in each language if available. We use open-access research archives such as arxivarxiv.org or HALhal.science. We also search for open-access research articles published under a Creative Commons license on Google Scholar using the same keyword in each language. We notice that research papers from natural sciences or technology are much less frequent in non-English languages as most researchers in those areas publish their work in English.
Literature: We collect sentences from different types of literature (Novels, History, Biographies, Children’s Stories) using books that are in the public domain. For English, French, and Russian, we use Project Gutenberggutenberg.org that archives old books for which U.S. copyright has expired. For Arabic, we use Hindawi Bookshindawi.org which provide free Arabic books in many genres and topics. For Hindi, the law in India states that the copyright terms of books end 60 years after the death of an author and comes under the public domainhttps://copyright.gov.in/Documents/handbook.html. Similar laws for most countries of the world are present with varying number of yearsen.wikipedia.org/wiki/List_of_countries%27_copyright_lengths. We thus manually search for books in Hindi whose copyrights have expired according to these lengths. For example, we used Hindi novels by Premchand, Sarat Chandra Chattopadhyay, Rabindranath Tagore and Devaki Nandan Khatri.
Textbooks: Textbooks are obtained from the Open Textbook Libraryopen.umn.edu/opentextbooks/books for English and Hindawi Books for Arabic which provide openly licensed textbooks. For Hindi textbooks, we use publicly available school textbooks from the National Council of Educational Research and Training in India ncert.nic.in/ which provides books at various high-school levels and in different subjects. No similar openly available resource was found for French and Russian.
Legal: We identify multiple governmental type of documents that we group under the "legal" domain, which include:
Constitutions: We sample sentences from the U.S. constitution for English, the Lebanese constitution for Arabic, the Indian constitution for Hindi, French constitution for French, and Russian constitution for Russian. Judicial Rulings: We used recent public decisions by law courts, such as the Supreme Court in the US law.cornell.edu/supremecourt/text, to collect sentences from judicial rulings, in addition to using legal datasets with such content Kapoor et al. (2022).
United Nations Parliament: We collect samples from the United Nations (UN) Parallel Corpus Ziemski et al. (2016) which contains official records and parliamentary documents of the UN. The corpus is available all languages we consider except for Hindi since it is not considered one of the official languages of the UN.
User Reviews: User text reviews for products, movies, books, hotels, and restaurants, are sampled from open-source datasets in each language when available. Most these datasets are used in sentiment analysis research.
Dialogue: Conversational text data is collected from three different types of open-source dialogue datasets: Open-domain dialogue datasets which focus on open-ended general conversation Naous et al. (2021); Li et al. (2017); Zhang et al. (2022), Task-oriented datasets that are design to train human-assistance or customer support dialogue models van der Goot et al. (2021); Malviya et al. (2021), and Negotiation dialogues that are used in developing automated sales dialogue agents with negotiation capabilities He et al. (2018).
Finance: We leverage the Financial Phrasebank dataset Malo et al. (2014) which provides English sentences with financial references and content collected from finance-focused news, and the CoFiF corpus Daudert and Ahmadi (2019) which provides financial reports in French.
Forums: We collect text from several online forums. These include: Reddit: Reddit is a popular platform where online communities discuss common interests and passions. We used the latest version of the Reddit dump available at the time of this study to sample user posts. We filtered posts for language using the fasttext language identification model with a confidence > 0.9. NSFW and Over 18 content were automatically filtered before sampling. Further, any sampled sentence that still contained sexual or offensive content was manually removed.
QA Websites: We collected questions and answers from QA websites using publicly available datasets for Question Answering research Nakov et al. (2016); Quora.com (2017); Howard et al. (2021); d’Hoffschmidt et al. (2020); Efimov et al. (2020).
StackOverflow: Sentences were collected from the StackOverflow NER dataset Tabassum et al. (2020) which contains user posts that describe what the user is trying to accomplish, a problem they are facing, or questions to seek advice from the community.
Social Media: We sample tweets from the the Stanceosaurus dataset Zheng et al. (2022) which provides thousands of tweets in English, Arabic, and Hindi that discuss recent region-specific rumors. Tweets that include offensive or hate speech were manually omitted.
Policies: We group under "Policies" several type of documents that delineate plans of what to do in a particular situation. This includes text extracted from: freely available contract templates for apartment/house leasing and job employment, Special Olympics rules which are available in multiple languages among which are but not in Hindi, and online codes of conduct of different organizations that we identify.
Guides: Several domains that aim at providing instructions to the reader are grouped under "Guides". We extract data from Samsung Smartphones User Manuals which are available in a variety of languages. Another source is Online Tutorials which we collect from WikiHow that provides how-to articles in multiple languages. We also manually collect Recipe Instructions from multiple online cooking resources for each language. Additionally, we collect Code Documentation sentences from documentation of different functions of the Matlab softwaremathworks.com.
Captions: We collect four different types of captions: image and video captions from various public datasets used in automatic captioning research, movie subtitles from the OpenSubtitles Lison and Tiedemann (2016) dataset used in machine translation research, and YouTube captions that we manually collect from video released under a Creative Commons license. While high-quality YouTube captions are easy to find for English, we could not find any high-quality YouTube captions for non-English languages.
Medical Text: We use clinical reports written by medical professionals from the i2b2/VA dataset Uzuner et al. (2011). We could not find similar high-quality medical resources for non-English languages.
Dictionaries: We manually collect sentence examples from Arabic and English dictionaries using words that have appeared in the Word of the Day. No similar resource under a Creative Commons license was found for Hindi, French, and Russian.
Entertainment: We use Humour detection datasets to collect jokes for Arabic Al-Khalifa et al. (2022) and English Weller and Seppi (2019). We manually collected jokes for Hindi.
Speech: Two types of sources for speech data are used: publicly available presidential speeches that are usually posted on governmental websites. We used speeches by the United States President that are posted on the department of state’s website. These speeches are also professionally translated to Arabic. We also collect sentences from TED Talk transcriptions, which are professionally translated from English to multiple languages.
Statements: Two different types of standalone sentences that we group under "statements" were identified which are: Rumours, and quotes. We collect rumours in Arabic, English, and Hindi from the Stanceosaurus dataset Zheng et al. (2022) used in misinformation detection. The rumours/claims are collected from various fact-checking websites in the Arab World, India, and the U.S. We also manually collected quotes in the three languages from various online resources. We did not collect mere translations of famous English quotes to other languages but focused on quotes by old scholars and thinkers of the Arab World, France, Russia and India for more cultural representation.
Poetry: Poetry lines are extracted from English, Arabic, and Hindi poems, some of which date back several centuries ago. To have culture specific samples, we focus on non-English poems from original Arab, French, Indian, and Russian authors, and not poems translated from English.
Letters: English letters were collected from online archives of historic letters. No high-quality authentic letters were found in Arabic or Hindi.
A.2 Domain Distribution
Table 5 shows the distribution of the domains in each readability level for each language. Basic readability levels (A1, A2) mostly contains sentences from domains that have text that is straightforward to read and contains day-to-day vocabulary such as Captions, Dialogue, User Reviews, User Guides. Intermediate readability levels (B1, B2) largely contain sentences from domains that present factual content such as books, Wikipedia articles, policy documents, news articles, etc. Proficient levels (C1, C2) contain domains that are scientific and technical such as finance, medical, legal documents, or highly literary text such as Arabic Poetry.
A.3 Examples
Example sentences from various domains are shown in Table 11 for English, Table 12 for Arabic, and Figure 14 for Hindi.
Appendix B Descriptions of CEFR Levels
The CEFR levels descriptions are provided in Table 6. Each level is described by specific reading capabilities which we used to familiarize annotators with the intuition behind the scale being used prior to labeling.
Appendix C Traditional Metrics
ARI and FKGL are statistical formulas based on text features including number words, characters, syllables.
ARI is a measure that aims at approximating the grade level needed by an individual to understand a text. It is computed as follows:
Flesch-Kincaid Grade Level (FKGL).
FKGL also aims at predicting the grade level, but unlike ARI, considers the total number of syllables in the text. It iss computed as follows:
Open Source Metric for Measuring Arabic Narratives (OSMAN).
OSMAN is computed according to the following formula:
where A is the number of words, B is the number of sentences, C is the number of words with more than 5 letters, D is the number of syllables, G is the number of words with more than four syllabus, and H is the number of "Faseeh" words, which contain any of the letters (\setcodeutf8\<ظ، ذ، ؤ، ئ، ء>) or end with (\setcodeutf8\<ون، وا>).
Appendix D Experimental Details
The details of the pre-trained language models used in our experiments are provided in Table 7, including the number of parameters and pre-training data sources. The majority of models have been pre-trained using CommonCrawl data. ARBERT Abdul-Mageed et al. (2021) and AraT5 Elmadany et al. (2022) are the models trained on the biggest number of sources.
D.2 Corpus Split
The train/validation/test split statistics of ReadMe++ are shown in Table 8 for each language. Those splits are obtained based on taking a 80%/10%/10% split for train/validation/test per domain, ensuring all domains are covered in each split.
D.3 Few-shot Prompt
The prompt used for Llama-7B is provided in Table 9. The prompt contains 5 primary features: The task description, definition of readability, example CEFR levels, example sentences with readability scores, and finally the new sentence for evaluation. When investigating the importance of the few-shot demonstrations we modified how we sampled the few-shot examples from the training set, however the prompt scaffolding remained the same.
Appendix E Additional Results
The F1 scores obtained by the fine-tuned models are shown in Figure 10. We also report the Spearman Correlation () as an additional correlation measure in Figure 11. The same trends for models observed in §4.4 hold for other metrics.
E.2 Domain Correlation
To explore the utility of the large data diversity in ReadMe++, we investigate the performance of models trained on both ReadMe++ and CEFR-SP across several specific domains. We train XLMRL using the publicly available Wikipedia splits of CEFR-SP (1 data source) compared to the public data from ReadMe++ (112 data sources) The correlation of model predictions with human annotated labels are shown for 21 different textual domains in Figure 12. In 18 out of the 21 domains, the model trained on ReadMe++ clearly outperforms the model trained on CEFR-SP underscoring the importance of data diversity in fine-tuning language models for readability assessment.
E.3 Zero-shot Cross Lingual Transfer
The zero-shot cross lingual results for several multilingual models are shown in Table 10. Similar to what is observed in §4.5, fine-tuning on ReadMe++ leads to significantly better cross-lingual transfer to 6 different target languages compared to fine-tuning on previous datasets. The improvement and trend is consistent across various multilingual models.
E.4 Effect of Context
We study the effect of providing models with context during training, which consists of up to three sentences that precede a sentence lying within a paragraph, on performance in the supervised setting. We prepend the context to the input sentence when available and separate them with a [SEP] token. Figure 13 shows the results with and without the addition of context when available.
Appendix F Annotation Interface
Figures 17 and 18 show screenshots of our developed annotation interface for English sentences, where annotators perform a rank-and-rate approach to assign readability scores to 5 sentences in each batch. Annotators are asked to first rank sentences which they can do by simply dragging them. They are then asked to choose a rating for each sentence from a drop-down list. For each sentence, we provide the option to show its context, which shows the sentence in the paragraph to which it belongs. Figures 19 and 20 show screenshots of the interface for Arabic and Hindi respectively. An additional button to mark transliterations is added.
Appendix G License and Use Terms
We provide in Tables 16, 17, and 18 the license or usage term for each data source used in the creation of the corpus as follows:
License: exact license under which data is available (CC BY 4.0 or other).
Public Domain: data available in the public domain.
Personal/Non-Commercial: source grants usage permission of data for personal/non-commercial purposes.
(—): denotes that data needs to be requested from authors.