Does Writing with Language Models Reduce Content Diversity?

Vishakh Padmakumar, He He

Introduction

Large language models (LLMs) are rapidly changing how people create content (Lee et al., 2022a; Mirowski et al., 2023). While LLM-based writing assistants have the potential to improve writing quality and increase author productivity, they also introduce an algorithmic monoculture (Kleinberg and Raghavan, 2021). As millions of users rely on the same underlying model to produce text, there is a potential risk of shifting the content towards the mode—good but not colorful writing. In this work, we aim to assess whether writing with LLMs unintentionally reduces content diversity.

Existing evidence has already hinted that LLMs may influence users’ opinions. Jakesch et al. (2023) and Bhat et al. (2023) find that user’s opinions expressed in their writing can be biased by controlling LLMs to promote a specific argument (e.g., social media is good/bad for society). Additionally, it has been shown that current LLMs do not equally represent views from various demographic groups (Santurkar et al., 2023; Durmus et al., 2023). These results suggest that writing with LLMs may limit the perspectives expressed in the writing. We hypothesize that incorporating model suggestions dilutes the writer’s unique voice, leading to different writers producing similar content, and this homogenization in turn reduces the overall diversity of content produced by many writers.

To test our hypotheses, we design a controlled experiment where users are asked to write an argumentative essay given a topic from the New York Times student opinion series following Lee et al. (2022a), e.g., “What are the most important things students learn at school?” We assign participants to three groups: a control group where participants write essays without model help, an LLM treatment group where participants write essays with a base language model (GPT3), and a feedback-tuned LLM treatment group where participants write essays with a language model finetuned with human feedback (Ouyang et al., 2022) (InstructGPT). An overview of the study is shown in Figure 1.

We hired 3838 writers from Upwork. For each group, we collected 100100 essays on 1010 topics. We then develop a set of metrics and measure the effect of LLMs on content diversity at both the individual level and the collective level:

Homogenization: Do users write more similarly to each other when writing with LLMs?

We find that essays written by the group using InstructGPT exhibit a higher degree of homogenization compared to the control group and the GPT3 group (Section 4). In particular, by matching model-contributed text to the summarized key points of each essay, we find that the model contributes around 40%40\% of an essay’s key points on average, which results in increased homogenization (Section 6).

Diversity: Does writing with LLMs reduce the diversity of content produced by a group of users?

We find that the set of essays written with InstructGPT does not only have lower lexical diversity, but also exhibits lower diversity in terms of the key points they present compared to the other two groups (Section 5). This is manifested through an increase in the repetition of phrases (higher-order nn-grams) introduced by the model (Section 5.2) and a reduction in the number of unique key points in the essays.

Interestingly, while writing with the human feedback-tuned LLM (InstructGPT) incurs increased homogenization and reduced diversity, we did not observe a statistically significant effect from writing with the base LLM (GPT3) despite a similar model contribution (Section 3). Further analysis shows that InstructGPT tends to provide less diverse suggestions and introduce less diverse text compared to GPT3. This finding echoes prior results showing reduced diversity after reinforcement learning from human feedback (Bai et al., 2022; Song et al., 2023).

As LLMs become more tightly integrated into text editing applications, our findings highlight a new axis of evaluation in interactive settings. Reduced content diversity is not only detrimental to personal expression and creativity, but also risks creating a feedback loop as future models trained on the homogenized content further propagate this pattern (Taori and Hashimoto, 2023). To facilitate future research on understanding the impact of LLMs in co-writing, we release essays from all three groups with keystroke-level information along with model suggestions recorded by the CoAuthor platform (Lee et al., 2022a).Code and data available at https://github.com/vishakhpk/hai-diversity.

Data Collection

Our approach to studying the impact of LLMs on content diversity is to conduct controlled experiments. Specifically, we control the set of writers and the topics of the essays. We then compare essays written with and without model help.

We consider the task of argumentative essay writing using topics from the New York Times Student Opinion series following Lee et al. (2022a). Specifically, users are given a topic such as “What are the most important things students learn at school?” and are asked to write an essay expressing their opinions in around 300300 words. We choose this task because the topics are sufficiently open-ended to allow for diverse responses while remaining accessible to users with all backgrounds. The instructions and the complete list of 1010 prompts used in the experiment are provided in Section A.3.

We adopt the CoAuthor platform developed by Lee et al. (2022a). When the user hits the Tab key, we obtain 55 text continuations from a backend model and present them in a dropdown menu. Figure 8 in Appendix A shows a screenshot of the interface as the user requests model suggestions. The user then has the option to either accept and edit one of the suggestions or reject all suggestions and continue with their writing process. We ask the users to request suggestions a minimum of 55 times per essay but do not require them to accept any of the suggestions. The entire writing process, including all keystrokes, suggestions, and user actions, is recorded by the platform. This allows us to differentiate the parts of the essay introduced by the users and models at the character level for subsequent analysis.

2 Experiment Setup

We recruit participants from Upwork who are fluent in English and have prior experience in writing or copyediting on the platform. We provide further details on participant recruitment and remuneration in Section A.1. We then collect essays from three settings: a control group, where participants write essays without any model help; an LLM treatment group, where participants write with the help of a base language model; and a feedback-tuned LLM treatment group, where participants write with a language model finetuned on human feedback. We distinguish the two types of language models because prior work has shown that finetuning language models on human feedback tends to decrease the entropy of the output distribution Bai et al. (2022) and thus may decrease text diversity in general. We henceforth refer to the three settings as Solo, GPT3, and InstructGPT respectively.

We obtain model continuations from the OpenAI API using davinci as the base language model and text-da-vinci-003 as the feedback-tuned language model. To avoid reduced diversity due to decoding, we use a decoding strategy biased towards higher entropy. Specifically, we sample continuations from both models with a temperature of 0.90.9 and a frequency penalty of 0.50.5 (detailed parameters listed in Section A.2), following the “high randomness” setting from Lee et al. (2022a).

Each participant is asked to complete a session consisting of three essay writing tasks—one in each of the three settings and each on a different topic. In each session, the order of the three settings and the assignment of topics are randomized. Each essay takes 1515 minutes on average and a session is usually completed under one hour. In total, we obtained 1010 essays on each of the 1010 topics for each setting, resulting in 300300 essays in total.

How much do users engage with the model?

Before diving into the analysis of content diversity, we first examine whether the user and the language models form an effective collaboration by examining the usage statistics and model-contributed text.

We find that users actively query the model and use model suggestions during writing. As shown in Table 1, users query the model around 9 times on average for each essay and accept about 70%70\% of them (note that they are asked to query at least 5 times). Since users may further edit the suggestions after accepting them, we want to know whether the accepted suggestions are retained in the final essay.Participants are not obligated to accept any suggestions from the model. We ask that they query the model for suggestions at least 55 times when writing an essay, accepting these only when appropriate (Section A.3). The interface used for data collection records keystroke-level information about whether each character was introduced by the model or the user. Thus, we report the average percentage of characters introduced by the model in each essay. We observe that the model contributes roughly 35%35\% of the characters in each essay, suggesting that they are reasonably helpful to the writing task.

Finally, we are curious whether the users find InstructGPT to be more helpful than GPT3 given its additional feedback tuning. However, the two language models appear to be equally helpful to users. We do not observe a significant difference between their usage statistics: an independent samples tt-test on the number of queries and acceptance rates shows no significance at the 5%5\% level.

Model contribution to key ideas.

The usage statistics tell us that LLMs contribute to a fair amount of the text written. However, are they contributing key arguments or merely supporting the elaboration of a point? To quantify model contribution to the main ideas, we need a way to represent the important content in an essay. Motivated by recent encouraging results of using LLMs for zero-shot summarization (Goyal et al., 2022), we summarize each essay into a list of key points by prompting gpt-3.5-turbo.We select gpt-3.5-turbo because of its high performance on summarization on the HELM leaderboard Liang et al. (2022). An example is shown in Table 2. On average, each essay is summarized into 6.866.86 key points with a standard deviation of 1.531.53.There is no statistically significant difference between the average number of key points generated for essays from any pair of groups based on independent samples tt-tests at the 5% level.

Given a list of key points for each essay, we then estimate the fraction of key points written by the model and the user. Specifically, we align each key point to a sentence in the essay with which it has the highest overlap, as measured by Rouge-L Lin (2004). If more than half of the sentence’s characters are model-generated based on the recorded keystroke data, the key point is attributed to the model; otherwise, it is attributed to the user. We show the fraction of key points written by the model for the two groups of co-written essays in Figure 2.

We observe that the median fraction for both GPT3 and InstructGPT is around 0.40.4, indicating that models make a significant contribution to the key content of the essay. However, there is also a large variation among users, with some essays having none or all of the key points attributed to the model, suggesting diverse levels of reliance on the model.

Having observed the significant contribution of models in co-written essays, we hypothesize that this model intervention would result in the users producing more similar content. Next, we test this hypothesis in Section 4.

Does writing with LLMs result in more similar essays?

In the previous section, we have seen that models contribute a substantial portion of the essay’s content in terms of both raw characters and key points. Does this lead to more similar essays, or do the essays retain the individual styles of the users? In this section, we measure the homogenization effect of co-writing with LLMs.

We first define the homogenization of a single essay as its average pairwise similarity to all other essays written on the same topic. Let DtD_{t} denote the set of essays on topic tt. The homogenization score of an essay dd (written in response to topic tt) is

where ∣⋅∣|\cdot| denotes the size of the set. Correspondingly, we define corpus homogenization as the average homogenization score of all essays. We use two metrics to compute the similarity between two documents: Rouge-L, an overlap-based metric; and BertScore Zhang et al. (2020), an embedding-based metric.We use the microsoft/deberta-xlarge-mnli model, the highest ranked model in the BertScore repository according to Pearson correlation with human ratings of similarity. Both homogenization metrics range from to 11 with a higher score indicating more similar content.

We measure homogenization at both the essay level and the key point level. For key point homogenization, we represent each essay as a concatenation of its key points (Section 3) before calculating its homogenization score.

2 Results

In Table 3, we observe higher corpus homogenization on essays written with InstructGPT than essays from the other two groups, with p<0.05p<0.05 using an independent samples tt-test. This result holds across both similarity metrics as well as when the essays are compared at the essay level (Table 10 in Appendix C). Figure 3 shows the homogenization scores across topics at the key point level using Rouge-L as the similarity metric.Figure 9 and Figure 10 in Appendix C contain similar plots of the homogenization scores at the essay level and key point level using both Rouge-L and BertScore. We observe that the InstructGPT group has a higher median homogenization than other groups on 77 out of 1010 topics. While the effect of model intervention varies across topics, the trend of increased homogenization persists.

Writing with GPT3 does not increase homogenization.

Although GPT3 and InstructGPT contribute a similar amount of text and key points to the co-written essays (Section 3), unlike InstructGPT, GPT3 does not appear to increase the similarity between essays written by different users. In particular, there is no statistically significant difference between the corpus homogenization score of GPT3 and that of Solo at the 5% significance level using an independent samples tt-test. This discrepancy between InstructGPT and GPT3 suggests that the impact of LLMs on content homogenization is not uniform. While InstructGPT often outperforms GPT3 on benchmarks across a wide range of metrics Liang et al. (2022), our results highlight that this performance boost may come at the cost of creating more homogeneous content in interactive settings. As language models are increasingly deployed online, documenting these limitations from a human-centered perspective will inform future work on adapting models with human feedback Casper et al. (2023).

Does writing with LLMs reduce the overall diversity?

In the previous section, we have seen that writing with the feedback-tuned model leads to different users writing similar essays on the same topic. A direct consequence of this homogenization effect is that it may reduce the overall diversity of the collection of writings from many authors. In this section, we test whether writing with LLMs reduce the diversity of the produced content.

To measure the diversity of a collection of essays, we represent the corpus as a bag of information units (e.g., nn-grams), and compute the number of unique information units divided by the total number of information units, such as the type-token ratio. Intuitively, if the essays all repeat each other (least diverse), the entire collection can be represented as a single essay (small ratio).

Specifically, we consider two types of information units: nn-grams and key points. Computing the fraction of unique nn-grams is straightforward and it has been used in prior work to measure lexical diversity (Li et al., 2023; Meister et al., 2023).

For key point diversity, we first represent each essay as a set of key points (as described in Section 3) and aggregate the key points from all essays. To compute the number of unique key points, we perform agglomerative clustering and consider all key points in one cluster to be equivalent. We use the implementation of agglomerative clustering from Scikit-learn Pedregosa et al. (2011) using the complete linkage criterion.Complete linkage measures the distance between two clusters by the maximum distance between any pair of items from them. We convert Rouge-L and BertScore to distance metrics for the clustering algorithm by subtracting the similarity score from 11 (which ranges from to 11).

2 Results

In Table 4, we show the diversity measured by nn-grams (where n=1,2,3,4,5n=1,2,3,4,5) for essays written in all three groups. We see that the InstructGPT group consistently exhibits lower diversity across all values of nn, whereas GPT3 does not significantly decrease the diversity score. The reduced lexical diversity showcases a disadvantage of co-writing with LLMs, particularly in creative writing genres like stories and poems, as it may wash away unique expressions on the long tail of human writings.

Writing with InstructGPT reduces key point diversity.

In Table 5, we report the diversity scores measured by the fraction of unique key points. The clustering process involves a threshold that determines the maximum distance above which two clusters will not be merged. Therefore, we report results for different thresholds for both distance metrics.A higher distance threshold indicates that key points having lower similarity can be clustered together. As a result, at higher thresholds, the diversity becomes lower for all groups because more key points are considered equivalent.

We find that the InstructGPT group consistently exhibits lower diversity across both similarity metrics and all thresholds. In addition, the diversity of the InstructGPT group is lower than both the Solo and InstructGPT groups by a statistically significant margin (p<0.05p<0.05) by a pairwise permutation test detailed in Section B.1. Perhaps more concerning than reduced lexical diversity, the decrease in key point diversity indicates that writing with LLMs may have a more profound social impact in the long run. As collaborative writing becomes more prevalent, care should be taken to ensure that perspectives from underrepresented groups are not lost in the writing process.

Essays written with InstructGPT repeat higher-order n𝑛n-grams more frequently.

To further analyze how diversity is changed, we plot the distribution of unigrams and 55-grams in each group of essays in Figure 4 (truncated to the most frequent 5050 nn-grams from each group).Each line in Figure 4 corresponds to the 5050 most frequently repeated nn-grams in a specific group, which vary across groups. Figure 11 in Appendix C shows similar plots for 11 to 55-grams. In co-writing groups, particularly in InstructGPT, a notable concentration of probability mass in the head of the distribution for 55-grams suggests a higher degree of repetition and lower text diversity in the co-writing process. We verify the significance of the difference in 55-gram frequencies between InstructGPT and Solo via a Chi-square test (p<0.05p<0.05).We detail the process in Section B.2. We observe a significant difference in nn-gram usage for 44- and 55-grams.

To see what kind of phrases the model repeats, we list the ten most frequent 55-grams in Solo and InstructGPT in Table 6. We find that common 55-grams from InstructGPT often contain topical words overlapping with the essay prompt, whereas those from Solo are more often generic phrases. As users may seek suggestions for key ideas in the essays (Section 3), the similarity of the suggested content manifests in these common topical 55-grams which occur repetitively.

Essays written with InstructGPT are more compressible.

Lastly, we show that the reduction of diversity measured by linguistic units such as nn-grams and key points also correlates with less diversity in an information-theoretic sense. Specifically, we measure the compression ratio of essays written in the three settings using various lossless compression algorithms: LZMA, ZLIB, and GZIP, each being a variation of the Lempel-Ziv algorithm Ziv and Lempel (1977, 1978). These algorithms compress data by identifying repeated patterns and using dictionary mappings for encoding. To compute the compression ratio, we concatenate all essays from a setting into a single text file. We then run the compression algorithm on it and calculate the ratio of the compressed file size to the original file size. From Table 7, we see that the InstructGPT essays are consistently more compressible across all three compression algorithms, which is statistically significant with p<0.05p<0.05 according to permutation tests (Section B.1).

Why does writing with InstructGPT reduce diversity?

A clear trend we observe is that writing with InstructGPT results in a reduction of content diversity, whereas this is not the case with GPT3. Since users seem to engage equally with both models (Section 3), this raises the question of how writing with InstructGPT makes a difference. There are two potential contributors. First, InstructGPT may produce less diverse text, thus influencing the essay by directly contributing text. Second, users may write differently in the presence of assistance, e.g., following the model’s phrasing or adopting the model’s opinion (Jakesch et al., 2023). In this section, we aim to answer this question by disentangling the effect from the model and the user in collaborative settings.

We first confirm that generations from InstructGPT are less diverse than GPT3, which has also been observed in prior work. For example, in the technical report for GPT-4, OpenAI finds that the feedback-tuned model is less calibrated OpenAI (2023); Bai et al. (2022) find that finetuning leads to decreased entropy of the output distribution. In our setting, we measure the diversity of model generation by the pairwise similarity of the five generated continuations upon each user query, similar to self-BLEU (Zhu et al., 2018). Thus, higher similarity scores indicate lower diversity. Figure 5 shows the box plot of the similarity scores computed using Rouge-L. We observe that InstructGPT indeed generates less diverse text than GPT3 on average (0.200.20 vs 0.110.11), with significance at the 5% level using an independent samples tt-test.We also obtain the same conclusion using BertScore as the similarity metric (see Figure 13 in Appendix C). Since users incorporate text from both models at similar rates (Table 1), it is plausible that the less diverse suggestions from InstructGPT lead to a decrease in the diversity of the final co-written essay. Next, we examine the diversity of model-written and user-written text in the essays directly.

InstructGPT increases repetition of higher-order n𝑛n-grams while user-written text is unaffected.

The co-written essays consist of both model-written and user-written text. Which contributes more to the decrease in diversity? To separate these effects, we attribute each token in the essay to either the user or the model.We use the recorded character-level keystroke information (Section 2.1) from the interface to decide if a token was written by the user or the model. Typically all characters in a token are written by either the user or the model. In the rare case where there is a mixed contribution, we assign authorship based on the majority of characters. We then examine the distribution of 55-grams from user-written and model-written text.We ensure that each entire nn-gram here was written by the user or model in a single continuous span of text to ensure we do not count incoherent nn-grams. From Figure 6, we see that the 55-gram distributions of user-written text remain the same regardless of whether the user writes with a model (Solo vs. GPT3/InstructGPT) and which model they write with (GPT3 vs. InstructGPT). This suggests that the phrase-usage pattern of users is not affected by the presence of model assistance. Instead, we observe that the increased repetition of 55-grams observed in Section 5.2 is more model related as evidenced by the high probability mass on common 5-grams for text introduced by InstructGPT.Each line in Figure 4 corresponds to the 50 most frequently repeated nn-grams in that setup, which vary across the setups. Figure 12 in Appendix C shows similar distributions varying nn from 11 to 55. We see the increased repetition of nn-grams introduced by InstructGPT in 33 and 44-grams as well.

InstructGPT increases similarity between key points while user-written text is unaffected.

Finally, we disentangle the effect from the model and the user for the increase in homogenization observed in Figure 3. We attribute each key point to either the user or the model (same as in Section 3) and calculate the essay homogenization scores on user-contributed and model-contributed key points respectively. In Figure 7, we compare the user-contributed and model-contributed homogenization across the InstructGPT and GPT3 groups. We observe that, while the average homogenization score of user-contributed key points does not change between the two groups, the InstructGPT contributed key points have higher average homogenization score than GPT3. An independent samples tt-test confirms the significance of this result at the 5%5\% significance level. This suggests that user behavior with respect to content creation is largely unchanged when writing with different models, We can also attempt to compare the homogenization of user-contributed key points with and without model assistance. However, since user-contributed key points of co-written essays are a subset of the original key points. They are more different from each other (less homogenized) than the full set due to randomness in subset selection, thus user-contributed key points in co-writing settings are not comparable to those in the Solo setting. and the increased homogenization is mainly due to InstructGPT contributions.

Takeaway.

Overall, these results suggest that the user’s writing behavior with respect to lexical choice and argument generation is not affected when writing with model assistance. The reduction in diversity in collaborative writing is attributed more to text contributed directly by the InstructGPT model. This paints a relatively optimistic picture as the risk of forming a feedback loop through the mutual influence between the user and the model is potentially limited. However, it is important to note that our study focuses on single interactions between users and models, and the dynamics might change through repeated interactions over time.

Related Work

Collaborative writing predates LLMs where users are assisted by suggestions retrieved from a knowledge base Swanson and Gordon (2008, 2012); Chen et al. (2014); Roemmele and Gordon (2015) or generated by task-specific RNNs Ghazvininejad et al. (2017); Roemmele and Gordon (2018); Clark et al. (2018).

The early systems can produce relevant suggestions, but they often need to be revised before being incorporated into the draft. In contrast, LLMs offer suggestions that are often directly incorporated into user writing, either through continuations or response to instructions. This has led to a surge of studies demonstrating the effectiveness of LLMs as writing assistants Gero and Chilton (2019); Akoury et al. (2020); Ito et al. (2020); Yuan et al. (2022); Swanson et al. (2021); Chakrabarty et al. (2022); Ippolito et al. (2022); Mirowski et al. (2023). However, despite the improved productivity, we have shown that LLM-contributed text can also affect the writing in subtle and unintended ways.

Effect of human-AI co-writing.

More recent work aims to analyze user-model interactions to understand they affect the user and their writing process. Buschek et al. (2021) study how users respond to varying numbers of suggestions provided in email writing and reveal the tradeoff between suggestion utility and writing efficiency. Bhat et al. (2023) show that model suggestions can be incorporated even when users disagree with the suggestions and their writing plans may be altered as a result of the interaction. Lee et al. (2022a) collect a dataset of user interactions with GPT-3 as a collaborator, observing that the model nudges users to use more diverse vocabulary, increase writer productivity, but had an uneven effect on the feelings of ownership towards to written content. Our work contributes to this line of exploration by examining the content homogenization effects that arise from user interaction with LLMs.

Evaluation of interactive text generation.

Traditional evaluation focuses on measuring similarity to the reference text Papineni et al. (2002); Lin (2004); Lavie and Denkowski (2009); Sellam et al. (2020). In interactive settings, due to the absence of references, most metrics aim to measure the effectiveness of model assistance indirectly, including user ratings of model helpfulness Roemmele and Gordon (2018); Clark et al. (2018); Padmakumar and He (2022), the fraction of model suggestions retained in the final text Akoury et al. (2020); Chakrabarty et al. (2022), and the time taken to complete the writing Buschek et al. (2021). Lee et al. (2022b) aggregate these into a framework for evaluating collaborative writing. They also reported user ownership and enjoyment, which emphasizes new aspects such as retaining the user’s writing style and influence on the content. Our work is related to existing metrics that evaluate lexical diversity and text style based on token and POS nn-grams statistics Roemmele et al. (2017); See et al. (2019); Tevet and Berant (2021); Meister et al. (2023); Tevet and Berant (2021). We extend them to co-writing settings and additionally develop more content-specific diversity metrics.

Social impacts of LLMs.

Bommasani et al. (2021); Bender et al. (2021); Solaiman et al. (2023) extensively discuss the societal risks of large-scale adoption of LLMs in user-facing applications. Most relevant to our work is the potential rise of an algorithmic monoculture Kleinberg and Raghavan (2021); Creel and Hellman (2022); Bommasani et al. (2022). While this line of work measures homogenization of outcomes in algorithmic decision making, our work looks at the impact of co-writing with LLMs in content generation. More closely related is recent work in human-computer interaction measuring the impact of writing assistance. Hancock et al. (2020) discuss how AI-mediated communication could result in the homogenization of language. Arnold et al. (2020) show that providing users with a predictive keyboard results in image captions that are shorter in length and contain fewer rare words. Jakesch et al. (2023); Bhat et al. (2023) show that writing with a biased language model can change user opinions. Liao and Xiao (2023) call for the development of new evaluation methods to identify the socio-technical issues associated with increased use of general-purpose LLMs. Our work investigates one axis of such evaluation, namely content diversity, in co-writing settings.

Conclusion

In this work, we study how writing with LLMs impacts the diversity of content produced. Through a controlled study, we find that users writing with InstructGPT produce more similar content than those writing with GPT3 or without model help. This homogenization also results in a reduction of the overall diversity of content produced by many users. Furthermore, our analysis indicates that this effect is attributable to the less diverse text contributed by InstructGPT, while the user-contributed text remains largely consistent with and without model help.

These findings highlight a new axis on which to evaluate the impact of LLMs prior to their deployment. While adapting a model with human feedback leads to an improvement in instruction following, this may be accompanied by a reduction in content diversity. Combined with other recent findings that LLMs influence user opinions during the writing process Jakesch et al. (2023), this calls for a careful user-centered evaluation to ensure that these models do not suppress the voice of users in scenarios where personal expression is desired. Finally, our work also adds to the growing body of open problems on reinforcement learning with human feedback Casper et al. (2023). Learning from the diverse feedback of many users and personalizing model generations to individuals are challenging future directions.

Acknowledgements

We would like to thank Mina Lee, Nick Lourie, Naomi Saphra, Richard Pang, and Nitish Joshi for their input at various stages of the project. We would also like to thank Abby Rabinowitz and Alexander Landfair from the NYU Expository Writing Program for valuable discussions during the planning stages. We would finally like to acknowledge all the writers recruited for this project without whom this work would not have been possible. This work is supported by Open Philanthropy, AWS AI, the Samsung Advanced Institute of Technology (Next Generation Deep Learning: From Pattern Recognition to AI), and the National Science Foundation under Grant No. 1922658.

Limitations

Our interface provides suggestions to users in the form of continuations of the current text in the draft. Further investigation is needed to evaluate if the reduction in content diversity with feedback-tuned models can be mitigated with prompt engineering or richer forms of interaction (e.g., through a dialogue).

User selection.

The outcome of the experiments can be affected by the specific group of participants. We try to ensure a diverse user group and detail our recruitment procedures in Section A.1. However, it is unclear whether these results will generalize to other groups such as students learning to write or second language speakers, who have different goals and incentives for using writing assistants.

LLM access.

Our experiments are conducted using two limited-access models from OpenAI. While we believe they are representative of properties of current LLMs, it is possible that the other models may exhibit different behavior, especially given that the RLHF pipeline is highly customized.

References

Appendix A Task Details

We recruit a total of 3838 participants from Upwork, all of whom are fluent in English and have prior experience in writing or copyediting on the platform. Their prior experience varied from 55 to over 200200 previously completed tasks on the platform. Our participants are exclusively from the United States and encompassing diverse racial backgrounds with an approximately equal distribution between genders. Each participant is asked to complete a session consisting of three essay writing tasks—one in each of the three settings and each on a different topic. In each session, the order of the three settings and the assignment of topics are randomized. Each essay takes 1515 minutes on average and hence a session usually takes the participants under one hour. Compensation was provided through hourly contracts, prorated to \20perhour.Aftereachparticipantcompletedonesession,theiressayswerereviewedmanuallybytheauthorsofthisworkfortherelevancetotheessaypromptaswellascoherenceofwriting.Thisprocessdoesnotinvolvepolicingtheactualcontentoftheessaysandweencourageourparticipantstoexpresstheiropinionsindetail.Thiswasusedtofilteroutafewparticipantswhosesessionswerereassigned.Theremainingparticipantsweretheninvitedtocompletetwoadditionalsessions,eachwithdifferenttopics.Asaresult,eachoftheper hour. After each participant completed one session, their essays were reviewed manually by the authors of this work for the relevance to the essay prompt as well as coherence of writing. This process does not involve policing the actual content of the essays and we encourage our participants to express their opinions in detail. This was used to filter out a few participants whose sessions were reassigned. The remaining participants were then invited to complete two additional sessions, each with different topics. As a result, each of the38participantscompletedbetweenparticipants completed between1andand3$ sessions. In total, we obtain 10 essays on each of the 10 topics for each setting, resulting in 100 essays per setting.

A.2 Decoding Parameters on OpenAI API

To avoid reduced diversity due to decoding, we use a decoding strategy biased towards higher diversity. Specifically, we sample continuations from both models following the “high randomness” setting from Lee et al. (2022a). The decoding parameters we use are:

Engine: davinci for the base language model and text-da-vinci-003 for the feedback-tuned language model

A.3 User Instructions

To complete an essay assignment successfully, you would need to write a short 3-passage essay given the prompt presented to you. The essay should reflect your opinion of the assigned topic and you’re writing an argumentative piece for why you feel that way. The expected length of the essays is around 300 words (3-4 short paragraphs of 3-4 sentences each). Each piece should take you between 10 and 15 minutes (as calculated from prior experiments). When writing with AI help, you are required to obtain model suggestions at least 5 times (we encourage you to do this more if you find it helpful, more the better). Some example responses are provided below (though these were written by non-experts so I’m sure you can identify flaws and improve on the style)

Instructions in Detail

When you are completing a task, make sure you have a Session ID. If you don’t have this already, please send us a message on Upwork.

View the spreadsheet of assignments and navigate to your assigned session. Each session will correspond to three rows in the spreadsheet, each with an accompanying URL.

Each of these links corresponds to the three assignment essays you would have to complete in order to finish the task.

If you click one of these, you will see the text editor pop up which looks like this which shows the prompt for the assigned essay.

Out of the three assignments, please note that one is the essay that you will have to write without the assistance of the AI. Please complete the three assignments in the order provided.

When you write with the AI, you will have the option to hit the TAB key on your keyboard and obtain suggestions from the model in the form of a dropdown like this. You are free to continue the writing process as you normally would, however, we encourage you to make use of model help when you are looking for ideas when you are stuck or just want to take a look at some possible continuations :)

When you write with the AI, you need to request at least 5 suggestions from the model. You do NOT need to accept all of the suggestions, you are free to accept one then edit it, ask for suggestions again at the same time, or even just reject all of the model suggestions. We encourage you to use the model as much as possible :) The objective of the study is to understand the effect of the model, so more interaction is better.

Once you complete the essay, make sure you hit ‘Save’ at the bottom of the screen to obtain a verification code on completion. We ask that you complete the writing in one sitting. We aim to record the entire writing process so please refrain from copy-pasting any text into the editor.

A.4 Summarization into Key Points

Motivated by recent encouraging results of using LLMs for zero-shot summarization Goyal et al. (2022), we summarize each essay into a list of key points by prompting gpt-3.5-turbo. Table 9 contains a full example of an essay with the generated key points. We use a simple prompt, “Summarize this essay into a set of simple and distinct bullet points. Make sure that the bullet points cover all the information from the essay.”

Appendix B Significance Testing

To test for the significance of the results on diversity (Section 5.2), since we do not have multiple instances of essay corpora from each setting (Solo, GPT3 and InstructGPT), we employ three permutation tests between each pair of settings. A walkthrough example of the permutation test setup is as follows. We wish to evaluate the significance of the difference between the diversity scores of Solo and InstructGPT, as measured by the clustering of key points (Table 5) or by traditional lossless compression algorithms (Table 7). We first calculate the statistic (difference between the diversity scores) on the collected corpora of essays. We then take the union of Solo and InstructGPT essays, randomly partition this into two equal sets, and recalculate the statistic on these two permuted sets of essays. This process is then repeated for 10001000 different permutations. We then obtain the p-value of this two-tailed permutation test as the proportion of times the absolute value of the statistic calculated on the permuted data is greater than the statistic calculated on the observed data from the user study. This is then repeated for all pairs out of Solo, GPT3, and InstructGPT. Bold values in the results tables (Table 5, Table 7) indicate those instances where the p-value on the permutation test for a setup (InstructGPT) was significant at the 5% level over both other setups (GPT3 and Solo).

B.2 Chi-Square Test for Significance for n𝑛n-gram Distributions

To confirm the significance of the difference between the categorical distributions of nn-grams in the different setups, we employ a chi-square test on the count of occurrences. We evaluate if writing with InstructGPT results in a change of 55-gram usage in Figure 4 and Figure 6. To perform this test, we first identify and take the union of the 5050 most frequently occurring 55-grams in each setup. We then obtain the counts of occurrences of each of these 55-grams in both corpora and perform a chi-square test for significance on these frequencies. A p-value <0.05<0.05 results in a rejection of the null hypothesis and the observation of a significant difference in the categorical distributions. We perform these for all pairs out of Solo, GPT3, and InstructGPT and find that the repetition of nn-grams when writing with InstructGPT is more similar by a significant margin compared to both other setups (Figure 4).

Appendix C Additional Results

In addition to the results presented in Section 4.2, we also report the corpus homogenization scores at the key point and essay level in Table 10. Across both Rouge-L and BertScore, the corpus of essays written with InstructGPT exhibits higher corpus homogenization than both other groups by a statistically significant margin (p-value <0.05<0.05) on both levels. We also plot the homogenization scores of all the essays calculated at the raw essay and key point levels in Figure 9 and Figure 10. We choose to report the homogenization at the key point level mainly because string similarity metrics such as Rouge-L and BertScore are less reliable on longer text documents Sun et al. (2019); Gehrmann et al. (2023).

Essays written with InstructGPT repeat higher- order n-grams more frequently

Building on from the results in Section 5.2, we plot the nn-gram distributions of the raw essays varying nn from 11 to 55 in Figure 11. Here we truncate the distribution to the most common 50 nn-grams from each setup. While the distributions for 11, 22 and 33-grams are almost identical across the setups, on 44 and 55-grams we see a concentration of probability mass at the head of the distribution for the essays written with InstructGPT. The reduction in lexical diversity (Table 4) is manifested in this increased repetition of higher-order nn-grams.

InstructGPT presents less diverse suggestions to users than GPT3

To test the hypothesis that adapting a model with human feedback reduces the diversity of the presented suggestions, we calculate the average similarity of all pairs of the five suggestions presented by InstructGPT and GPT3 upon each user query. In addition to the similarity plotted with Rouge-L in Figure 5 we also plot the average similarity scores of suggestions computed using BertScore in Figure 13.

Higher correlation between AI written fraction of the document and homogenization, on InstructGPT than GPT3

To study the relationship between increased model intervention and homogenization, we plot the fraction of the essay written by the model as a function of the document homogenization with BertScore and Rouge-L in Figure 14. We observe a weak correlation when users write with InstructGPT, particularly in the case of BertScore.