The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text

Yanzhu Guo, Guokan Shang, Michalis Vazirgiannis, Chloé Clavel

Introduction

Scaling law reveals a predictable smooth increase in model performance as the amount of data, compute power, and model parameters are amplified in tandem (Ganguli et al., 2022). Even assuming that we can boost the other two ingredients indefinitely, the amount of data is limited. By one estimate, the world’s entire supply of high-quality text ranges up to 17 trillion tokens, with a 4-5% yearly growth rate (Villalobos et al., 2022). This includes all the world’s books, scientific papers, news articles, Wikipedia pages, available code, and the rest of filtered web content. Meta’s Llama 2, one of today’s leading LLMs, was trained on around 2 trillion tokens (Touvron et al., 2023). In other words, we might be approaching the exhaustion of the world’s entire stock of usable language training data, potentially within an order of magnitude.

Is it possible for LLMs to train on their self-generated samples, thereby offering a solution to the looming data shortage? In fact, whether intentionally or unintentionally, this would happen with the widespread recognition and usage of LLMs. Regarding pretraining data, which is often sourced from the Internet, a significant trend is occurring: an increasing volume of online content is either generated or assisted by models, and such content is nearly indistinguishable from data produced by humans (Uchendu et al., 2023). Consequently, the subsequent generations of models will inevitably be pretrained on deeply blended data. Regarding finetuning data, employing LLM-generated examples is already a widely adopted data augmentation approach in the NLP community. The work of self-instruct Wang et al. (2023) prompts language models to solicit synthetic multi-task instruction-tuning data in an iterative bootstrapping way, starting with a seed set of manually-written instructions. Concerning single-task training, Zhou et al. (2023) build a large-scale multi-scenario multi-domain dialogue summary corpus annotated by ChatGPT (Ouyang et al., 2022) to enhance their pretrained dialogue summarization model.

However, recent studies raise concerns that the above approach of training on predecessor-generated text—LLMs are trained on the data produced by previous models—is not a panacea without side effects, especially when conducted recursively over time. This would introduce a new set of challenging issues, a phenomenon described as model collapse (Shumailov et al., 2023; Alemohammad et al., 2023). On one hand, incorporating model-generated content in training may lead to irreversible flaws in the resulting models, where tails of the original distribution of genuine human content disappear. On the other hand, even when these models remain free of defects, they could converge to excessively uniform behaviours, with very small variance, due to the recursive sampling of only probable events over generations.

In this study, rather than focusing on shifts in task-solving performance, our primary interest lies in exploring changes in language usage caused by the degenerative learning process of recursively training language models on the data generated by their predecessors. Our work is motivated by and contributes to, answering the following two key research questions: First, how can linguistic diversity be quantified effectively? Second, does recursive training on predecessor-generated text result in a reduction of linguistic diversity in model outputs?

To address these questions, we first develop a comprehensive set of novel metrics aimed at assessing diverse aspects of linguistic variation, encompassing lexical, semantic, and syntactic diversity. Subsequently, we proceed to conduct a series of recursive finetuning experiments spanning three natural language generation tasks, each demanding varying levels of creativity: news summarization (Hasan et al., 2021), scientific abstract generation, and story generation (Fan et al., 2018). Our results indicate a notable trend: with the progression of recursive finetuning iterations, there is indeed a remarkable decrease in the diversity of the generated outputs. This observation highlights the significant impact that training on text generated by predecessors has on the linguistic diversity of LLMs.

Related Work

In this section, we explore two avenues of related work: current approaches to evaluate linguistic diversity and recent research on training with synthetic data generated by language models.

Efforts to evaluate language models predominantly concentrate on their performance in task-solving. While some studies extend their scope to include aspects like factual consistency, reasoning capability, and robustness (Chang et al., 2023), there is a notable lack of attention to linguistic diversity, which remains significantly understudied.

Furthermore, the existing studies that do address the diversity issue typically focus on lexical diversity alone. For example, in quantifying diversity, research on decoding strategies (Li et al., 2022b; Vijayakumar et al., 2018; Ippolito et al., 2019) usually considers the proportion between the number of unique nn-grams and total number of nn-grams in generated text, known as distinct-nn metric (Li et al., 2016). This very approach can also be found in the literature related to specific NLG tasks, especially those regarding creative text generation, such as poetry (Chakrabarty et al., 2022), lyric (Tian et al., 2023), and pun (Mittal et al., 2022) generation.

Alternatively, Zhang et al. (2021) propose to use Shannon entropy to quantify diversity. However, such approach is still calculated on the tokens (i.e., lexical level), demonstrating a strong correlation with distinct-nn. Zhu et al. (2018) introduce Self-BLEU which calculates the BLEU similarity score (Papineni et al., 2002) between different sentences of the same document, with higher Self-BLEU implying lower diversity. This metric is adopted as a proxy for diversity in evaluating the capability of LLMs in the context of producing content for disinformation operations (Liang et al., 2022). Nevertheless, the BLEU score is based on nn-gram overlap and thus also represents diversity solely from the lexical aspect.

Very few works study diversity beyond the lexical level. As an exploratory approach to quantify syntactic diversity, Clercq and Housen (2017) first manually annotate a small corpus of texts produced by second language learners for syntactic features such as syntactic length and clause types, whose variation is then viewed as a diversity index. Recently, Padmakumar and He (2023) bring up the semantic aspect of diversity and define the average pairwise BERTScore among a set of documents as the homogenization index. They also use ChatGPT to annotate key points on a small set of documents, counting the percentage of unique key points as content diversity.

Our work is the first to comprehensively evaluate text generation on all three aspects of linguistic diversity: lexical, syntactic and semantic. In addition to utilizing the well verified lexical diversity measures, we also propose novel metrics for automatically evaluating syntactic and semantic diversity on a large scale.

2 Training with Synthetic Text

Ever since the introduction of generative adversarial networks (Goodfellow et al., 2014), training new models with synthetic data produced by various generators has become a means of data augmentation (Li et al., 2022a), a practice that has been expanding to all modalities of machine learning research, including image, audio, and text.

However, the large-scale usage of this approach, particularly employing tremendous quantities of synthetic text to train generative models, is a more recent trend (Dai et al., 2023; Marwala et al., 2023). To name a few, the self-instruct study by Wang et al. (2023) guides a language model to iteratively generate synthetic multi-task instruct-tuning data, beginning with an initial set of manually-written instructions. Huang et al. (2022) demonstrate that LLMs are capable of self-improving their reasoning abilities, with generated high-quality answers for unlabeled questions, using chain-of-thought (Wei et al., 2022) and self-consistency (Wang et al., 2022) prompting techniques. Meanwhile, Xu et al. (2023) introduce a pipeline that autonomously generates a high-quality, multi-turn chat corpus by enabling ChatGPT to converse with itself, which is then used to enhance a LLaMA model.

As aforementioned in Section 1, concerns have been raised about such training methodology (Shumailov et al., 2023), especially considering the impact on the future LLMs, generation after generation. Training data (often crawled from the Internet) will be abundantly generated by previous models and LLMs will be recursively trained on predecessor-generated text. Studies show that this process will eventually lead to model collapse, causing performance degeneration, regardless of potential data filtering or refinement (Alemohammad et al., 2023). Our research is motivated by the same concept, yet our main focus is on investigating how linguistic diversity is affected by this degenerative learning process, rather than on the decline in performance.

Methodology

This section introduces our recursive training methodology and outlines the metrics employed to evaluate both performance and diversity.

Following the work of Shumailov et al. (2023), we simulate the process of recursively training LLMs on predecessor-generated text, under a finetuning setting. As illustrated in Figure 1, we begin with human-generated task-finetuning Data (0), which is used to train Base (1) model to create a task-specialized version, referred to as Model (1). After that, we use Model (1) to produce synthetic task-finetuning Data (1), which serves to train the next generation, Model (2), built upon Base (2) model. This procedure is repeated until we reach the nn-th generation. Due to finite sampling, we are only able to approximate the original distribution via fitting, and errors compound over time during this process.

For the sake of simplicity, we start from a new instance of the same base model across different generations, i.e., Base (1) = Base (2) = ,...,,..., = Base (nn). In addition, we only use Data (n−1n-1) to train Model (nn), however in a setting closer to the real-life scenario, we have access to the accumulated data ensemble of all predecessors, i.e., Data {\{(0), (1) ,...,,..., (n−1n-1)}\}. This simplification draws from the results of Shumailov et al. (2023), indicating that model collapse is unavoidable, even if the training involves the full ensemble of accumulated data or a sampled subset version, though the effect is somewhat attenuated.

In terms of finetuning tasks, we chose three distinct natural language generation tasks, each characterized by varying degrees of constraint, from the most restrictive to the least: news summarization, where summaries must closely align with the original content; scientific abstract generation, with some initial context provided, but room for creative expansion; and story generation, which allows for the most creativity and freedom in expression.

In the end, we conduct our linguistic diversity research with the fine-tuned Model {\{(1), (2) ,...,,..., (nn)}\} for each task, subjecting them to evaluation on the test set of corresponding task.

Perplexity

Our research primarily centers on linguistic diversity, yet we also require a reliable metric to verify that our fine-tuned models are well-aligned with the training data. Perplexity, a standard metric for assessing language modeling, evaluates a model’s level of "surprise" or "confusion" when encountering a given sequence of tokens. Models that more accurately mirror the training data’s distribution exhibit lower perplexity. While useful for model comparison, perplexity doesn’t fully reflect text quality (Meister et al., 2023). A low perplexity score suggests higher predictive precision, but texts can be grammatically sound and contextually coherent yet still score high in perplexity if they include unusual or creative language not included in the model’s training data (Basu et al., 2021). In our study, a model with lower perplexity isn’t deemed superior by default. Our aim is to ensure that perplexity remains within a reasonable limit, producing texts of sufficient quality for our linguistic diversity evaluation.

We approach the evaluation of linguistic diversity from three different perspectives: lexical diversity, semantic diversity and syntactic diversity.

Lexical diversity metrics are used to measure the variety of words used in a text, which is contended to mirror the extent of vocabulary possessed by a writer or speaker. We believe a degenerated LLM, who presumably has a smaller vocabulary, will use a narrower variety of lexical items than non-degenerated LLMs. We select different metrics operating at different levels of textual granularity: word, nn-gram, and sentence.

Type-Token Ratio (TTR) (Johnson, 1944; Templin, 1957), the most well-known metric, which is calculated as the number of unique words (types) tt divided by the number of running words (tokens) cc, i.e., TTR=\nicefractc\text{TTR}=\nicefrac{{t}}{{c}}. This metric was used to study the language development in child language research, a low value is probably indicative of a language-specific deficiency (Miller, 1981). The length of a text inherently skews vanilla TTR values, with longer texts generally yielding lower TTR scores due to an inexorably decreased occurrence of unique novel words (draw from a limited vocabulary) as the text lengthens (Richards, 1987). Following a common practice, we truncate all texts to fixed length before computing the TTR.

Distinct-n\mathbf{n} (Li et al., 2016), which is characterized by the proportion between the number of unique nn-grams and total number of nn-grams in tested text (Xing et al., 2017). This metric is originally introduced in the context of enhancing the response diversity of dialogue generation LLMs, which frequently produce safe and fluent but dull and uninformative generic responses at time (e.g., I don’t know) (Han et al., 2022). Similar to naive TTR, distinct-nn varies as a function of text length, so we report the results at fixed sizes for n=2n=2 and n=3n=3 (distinct-11 is equivalent to TTR).

Self-BLEU (Zhu et al., 2018), a recently developed method for evaluating the diversity of artificially generated text. This method assesses the similarity between one sentence and the rest in a group of generated sentences. It treats one sentence as the hypothesis and the others as references to calculate BLEU scores (Papineni et al., 2002). The final Self-BLEU score averages these BLEU scores across all generated sentences. We use a publicly available implementation (Alihosseini et al., 2019)https://github.com/Danial-Alh/fast-bleu. We report 1−Self-BLEU1-\text{Self-BLEU}, so a higher value reflects richer diversity of the generation (Palumbo et al., 2020).

1.2 Semantic diversity

According to recent studies (Tevet and Berant, 2021; Stasaski and Hearst, 2022), the above lexical-level evaluation metrics often fail to capture semantic diversity, since texts including similar words can have different semantics and texts with different words can have similar semantics (Yarats and Lewis, 2018). We tackle this problem by transforming sentences into semantically meaningful sentence embeddings using Sentence-BERT Reimers and Gurevych (2019). We quantify semantic diversity as the dispersion of sentence embeddings over the semantic space. The dispersion is measured by either the average pairwise cosine-distance of all embedding vectors (D_sem_p) or the mean cosine distance of each embedding vector to the centroid (D_sem_c).

1.3 Syntactic Diversity

The significance of syntactic diversity is often underestimated in natural language processing, despite its importance. For language learners (as well as language models), exposure to a wide range of syntactic structures is beneficial for developing a more comprehensive understanding of the language (Aggarwal et al., 2022). Moreover, a range of syntactic forms enhances expressiveness and subtlety in writing, influencing the style and tone of a text (Edwards and Bastiaanse, 1998). While linguistic and language acquisition research (Clercq and Housen, 2017) has explored this aspect, these studies typically rely on manual annotation of features, a process that can be costly and prone to human error.

Our research introduces the first automatic metric to quantify syntactic diversity. We use a neural parser (Qi et al., 2020) to construct dependency trees from sentences, following the Universal Dependencies formalism. These trees are then transformed into graph representations, with nodes denoting the words and edges indicating the dependency relationships between them. Subsequently, we employ the Weisfeiler-Lehman graph kernel (Shervashidze et al., 2011; Siglidis et al., 2020) to map these graphs into a vector space. This kernel, rooted in the Weisfeiler-Lehman isomorphism test, effectively positions graphs that are structurally alike closer to each other in the embedding space. To assess syntactic diversity, we calculate it similarly to semantic diversity, using either the average pairwise distance (D_syn_p) or the average centroid distance (D_syn_c) among all graph embeddings.

Experimental Setup

We conduct our experiments on three generative tasks, as introduced in Section 3.1, with decreasing degrees of constraint: abstractive news summarization, scientific abstract generation, and story generation. Separately for each task, we simulate 6 iterations of the recursive training chain, i.e., n=6n=6 on Figure 1. Following previous work Shumailov et al. (2023), we select OPT (Zhang et al., 2022) as our base model, and each iteration begins with a new instance of the base model. We use OPT-350M instead of OPT-125M to maintain the quality of generated text over iterations, avoiding excessive noise. Model (1) is fine-tuned on Data (0)—the training set of the finetuning task—which is human-authored. From Model (2) to Model (6), they are fine-tuned on synthetic Data (n−1n-1) generated by their predecessor Model (n−1n-1). We go through all of the original training examples in Data (0) to produce a comparable synthetic dataset of the same size. The models are finetuned for 5 epochs with AdamW optimizer (Loshchilov and Hutter, 2017) on a cluster of two NVIDIA RTX A6000 GPUs.

In the following, we explain each of the three tasks in detail.

For abstractive news summarization, we use the XL-SUM dataset (Hasan et al., 2021), one of the most recently proposed datasets for this task. In comparison with the other prominent news summarization datasets, XL-SUM is more abstractive than CNN/DailyMail (Hermann et al., 2015; Cheng and Lapata, 2016) and more factual than XSUM (Narayan et al., 2018; Guo et al., 2022). It is also larger in scale, consisting of 306522 samples in the training set and 11535 samples in the test set. The average length of the news articles is 386 tokens, while the average length of the summaries is 28 tokens. This is a generation task with “low entropy” since there is abundant context and the content is restricted.

Task2: Scientific Abstract Generation

For scientific abstract generation, we parse the bibliography database (BibTeX) file with abstracts of ACL Anthologyhttps://aclanthology.org in June, 2023. ACL Anthology hosts papers published at computational linguistics or natural language processing venues since 1965. We split the bibliography entries into the training and the test set, resulting in 40512 samples for train and 2132 samples for test. We use the title of the paper and the first sentence of the abstract as the prompt, asking the model to finish generating the rest of the abstract. The prompt (title + first sentence) is 42 tokens long on average, while the mean of the full abstract length is 145 tokens. This is a task of “medium entropy”, the provided title already lays out the general idea of the paper and the first sentence provides a fair amount of context.

Task3: Story Generation

For story generation, we use the WritingPrompts dataset (Fan et al., 2018). It is made up of human written stories paired with writing prompts from Reddit’s WritingPrompts forum. There are 545200 samples in the training set and 30276 samples in the test set. The writing prompts consists of 30 tokens on average and the resulting stories 389 tokens. The prompts are generally short and in most cases do not contain a plot (narrative structure), making this a “high-entropy” generation task with limited context.

We use a combination of nucleus sampling (pp) and temperature sampling (τ\tau) to achieve nuanced control over the language model’s outputs. We adapt the specific parameter to the characteristic of each task. For news summarization, we emphasize precision and set p=0.1p=0.1, τ=0.3\tau=0.3. For story generation, we care more about creativity and set p=0.9p=0.9, τ=0.7\tau=0.7. For scientific abstract generation, we want something in between and set p=0.5p=0.5, τ=0.5\tau=0.5. The max_new_tokens value is chosen according to the length of human-written references for each task: 50 for news summarization, 500 for story generation and 300 for scientific abstract generation.

Results

In Table 1, we display the perplexity and linguistic diversity metrics for texts generated across various iterations. Our findings indicate that the perplexity values fall within an acceptable range (Holtzman et al., 2020), suggesting that models effectively assimilate training data and generate texts of a quality viable for diversity analysis. A general decline in all linguistic diversity metrics underscores the pressing issue of diminishing linguistic diversity. We highlight some key observations in the following.

We deliberately select three language generation tasks of varying “entropy”, which is reflected by the amount and nature of given context, i.e., constraint. News summarization involves a lengthy context with the summary confined to a very limited space, whereas story generation is characterized by brief prompts and a vast array of potential narrative directions. In the highest "entropy" task, story generation, the gap in linguistic diversity between human-written and model-generated texts is the most pronounced, and the decline over iterations is the fastest. The significant gap between humans and models is expected, given that story generation demands substantial creativity, a domain where language models are known to fall short (Chakrabarty et al., 2023). The rapid decrease in diversity can also be explained by the creative nature of the task. Models initially learn from diverse original human-written stories but suffer greatly when later exposed solely to synthetic data, which already exhibits a notable loss in diversity.

In tasks like news summarization and scientific abstract generation, which have “lower entropy” compared to story generation, there is still a noticeable decrease in linguistic diversity over iterations. Consider the task of generating scientific abstracts: initially, the syntactic diversity in texts created by Model (1) shows an increase compared to those written by humans. This might be because scientific abstracts inherently possess less varied syntactic structures than the broader range of texts in the pre-training data of OPT-350M. However, as the iterations advance, the syntactic diversity scores of the texts produced by the model steadily decline, eventually dropping significantly below those of human-written abstracts. This trend might be partly attributed to catastrophic forgetting McCloskey and Cohen (1989). Additionally, while human-written abstracts may have limited syntactic diversity, their structure is markedly different from the pre-training data, thus introducing new learning elements for the model. In contrast, the synthetic data produced by Model (1), despite its marginally higher internal syntactic diversity, closely mirrors the model’s own training distribution. This lack of novel information leads to a subsequent reduction in variation.

We notice that syntactic diversity consistently decreases across all three tasks, comparable to the decline in lexical diversity and to a greater extent than in semantic diversity. While the reduction in lexical diversity is well-researched and somewhat anticipated, our study is the first to highlight the decrease in syntactic diversity. Syntax, although more implicit, is equally important as lexicons in maintaining linguistic richness. The significant yet often overlooked decline in syntactic diversity emphasizes the need for future natural language generation research to include syntactic diversity measurements alongside the commonly reported lexical diversity metrics.

Conclusion

Our study provides critical insights into the implications of recursively training LLMs on synthetic data generated by their predecessors. Through our innovative approach, focusing on novel linguistic diversity measures rather than traditional performance metrics, across various NLG tasks, we have uncovered a noticeable reduction in lexical, syntactic, and semantic diversity in LLM outputs over successive iterations of recursive training on synthetic text. These findings highlight a concerning trend: as LLMs increasingly rely on predecessor-generated text for training, there is a tangible risk of diminishing linguistic richness and variety in their outputs. Our research underscores the necessity for a more nuanced and forward-thinking approach in the development of LLMs, emphasizing the importance of preserving linguistic diversity alongside improving technical performance.

References