Is GPT-3 Text Indistinguishable from Human Text? Scarecrow: A Framework for Scrutinizing Machine Text

Yao Dou, Maxwell Forbes, Rik Koncel-Kedziorski, Noah A. Smith, Yejin Choi

Introduction

Clark et al. (2021) demonstrated the challenges of human evaluation in the era of GPT-3 Brown et al. (2020), as crowd workers are no longer able to reliably distinguish GPT-3’s generations from human-written text.

Or are they? In this paper, we propose a new framework for systematically scrutinizing machine text so that even crowd workers, despite the known challenges reported by recent literature, can successfully critique seemingly fluent generations. We not only quantify a measurable gap between machine text and human text, but reveal the distributions of specific categories of issues, and pinpoint their occurrences in text written by several sizes of language models as well as humans.

To achieve this, we develop Scarecrow, a methodology for eliciting categorical judgements of errors in machine-generated text from crowd workers. One goal in natural language generation (NLG) is to produce fluent outputs which can be read by laypeople. As such, we propose that important errors to address are those which are recognized by readers without NLP expertise. Our framework allows crowd workers to annotate problems in model outputs at the span level. A single such annotation is shown in Figure 1.

To make this possible, we establish a categorization of shortcomings commonly found in machine generated text (Table 1). This error schema covers a broad scope of problems as identified by experts, but has been honed according to what is salient to non-expert readers through several pilot rounds of crowd annotation without a fixed label set. The result is a framework that is usable by everyday people with minimal training, but covers the error phenomena found in real machine-generated text. Labeling spans of text using specific error types creates a picture of contemporary model generations with an unprecedented level of detail. In contrast to judging text holistically Celikyilmaz et al. (2021), insights from this method are specific and practical, as it measures exactly how and where problems arise.

We conduct a large-scale analysis of human-written and machine-generated text using Scarecrow, collecting 13k annotations of 1.3k paragraphs, amassing 41k spans labeled with error type, severity, and an explanation. Through this, we characterize in which ways GPT-3’s generations are better than those of previous models, and which aspects do not improve with increased data and parameters. We also provide a rigorous error analysis of text generated by several other contemporary language models, examining the impact of model size, training data, and decoding strategy.

We provide our detailed annotator training system and task interface so that future researchers may employ and refine them for error analyses of machine-generated text. We hope this will contribute to the standardization of NLG human evaluation Howcroft et al. (2020).

Key Findings

We perform a large-scale annotation of errors in English news text generated by five sources (four models and ground truth articles). We present Figures 2, 3, and 4 as summaries of our main results. As a reminder to readers, Grover Zellers et al. (2019) is the same model size and architecture as GPT-2 XL Radford et al. (2019), but trained in-domain (on news text). As such, our results cover three increasing model sizes (GPT-2 Small, XL, and GPT-3 Brown et al. (2020)), one change in domain (Grover), and ground-truth text (Human). For GPT-3, we also study a variety of decoding configurations (Figure 4).

The main quantity we measure (on yy-axes) is span coverage, which is the average portion of tokens that ends up covered by annotations of a particular error type. Since it is possible that multiple spans nest or overlap, there is no upper bound for this quantity. (See Figure 12 for a comparison of span coverage with other measurement alternatives.) Figure 2 measures span coverage for each type of span separately, Figure 3 stacks them, and Figure 4 removes non-error spans (reader issues) before adding them (as in Figure 3, but without showing the individual types).

1. Scaling pays off to improve [ Encyclopedic], [ Commonsense], and [ Incoherent] errors (Fig. 2). These error categories decrease with in-domain training (Grover) and larger model size (GPT-3). Human text still shows the fewest of these kinds of errors.

2. Scaling benefits plateau for [ Off-Prompt], [ Bad Math], and [ Grammar and Usage] errors (Fig. 2). These three error categories see a model plateau in error reduction when scaling to GPT-3. Of these error types, humans still commit fewer [ Off-Prompt] (more: §E.1) and [ Grammar and Usage] errors, but [ Bad Math] appears saturated for our domain.

3. [ Self-Contradiction] and [ Redundant] errors exhibit more complex scaling behavior (Fig. 2). We roughly categorize these trends as rising and falling: increasing for medium or large-scale models, but dropping for human-authored text. Text generated by GPT-2 Small is so often incoherent that there is little possibility for [ Self-Contradiction] (more: §E.2), and the increase in [ Redundant] errors varies based on how errors are counted (more: §E.3).

4. Human-authored text produces the most reader issues (Figs. 2 and 3). The [ Needs Google] and [ Technical Jargon] span categories both have a humans highest trend, and both fall under reader issues: problems that are not necessarily errors, but that still prevent full comprehension or factual verification of the text (more: §E.4).

Furthermore, human-authored text is not free from error annotations (Figure 3). This can serve either as a control for baseline error rates (more: §E.6), or as a mechanism for critiquing human writing.

5. Decoding hyperparameters have a huge impact (Figure 4). For the previous findings, we fix the sampling configuration for all models to an apples-to-apples setup for fair comparison: top-pp = 0.96, (softmax) temperature = 1, and no frequency penalty (i.e., word repetition penalty; defined precisely in §5.2, Equation 1). To study the effects of these decoding settings, we annotate text generated by GPT-3 using a variety of values for top-pp and temperature, both with and without a frequency penalty.

To our surprise, the decoding hyperparameters considerably affected error rates (more: §E.5). As seen in Figure 4, the worst sampling procedure for GPT-3 (argmax sampling with no frequency penalty) performed even worse than GPT-2 XL. But the best sampling procedure (surprisingly, also argmax sampling, but with a frequency penalty) produced text with as few apparent Scarecrow error spans as those authored by humans (more: §E.6).

All of these findings are discussed in more detail in Appendix E.

Evaluation of Natural Language Generation

We make our study in the area of open-ended natural language generation, a loose term for generating longer texts with an increased level of creative freedom. The common factor in all open-ended generation tasks such as story, blog, and dialog generation is the wide and diverse nature of target outputs. Lexically and even semantically dissimilar responses to the same prompt could be equally valid. For example, a model prompted with the blog title “Recipes for success this Holiday season” could describe how to roast a turkey or strategies for dealing with the stresses of holiday travel.

This allowable variation poses a particular difficulty for the evaluation of generation systems. Traditionally, text generation quality for tasks like machine translation or graph-to-text generation has been measured by word overlap with human-authored references (Papineni et al., 2002; Lin, 2004). Though measures like BLEU allow for multiple references, they break down when the space of allowable outputs is large, as in open-ended generation. Recently introduced metrics seek to remedy this problem (Hashimoto et al., 2019; Pillutla et al., 2021), but the gold standard for evaluating generated text is still human judgment.

However, current approaches to eliciting human judgement of generated text often do not provide detailed insight into where models are making progress, where they are failing, and the scope of these failures. A/B-style testing allows for directly comparing one system against others Clark and Smith (2021), but can only express relative improvements. Simple Likert scale judgements can assess text quality, but do not explain why a generated text receives a given rating, or which segment of the text is problematic. Insights into model failures often come instead from a small scale expert analysis of outputs. However, these “error analyses,” once a staple of NLP research, have become less common in recent years, perhaps due to their small size and high variance.

A hypothesis of the current work is that a well designed error analysis annotation framework could be used by crowdworkers to annotate large amounts of text, thereby providing detailed information about model progress and failures as well as actionable directions for future research. Such a framework would be easy to learn, reusable, and independent of particular models or experimental conditions. In what follows, we outline the details of such a method.

Scarecrow Annotation Methodology

This section describes the high-level annotation methodology for Scarecrow.

Our annotations consider two segments of text: a one-sentence prompt, and a one-paragraph generation. The prompt is human-written. It provides both starting tokens for model generation, as well as context for humans to evaluate whether a model is able to stay on-prompt—both topically and factually. Annotators know that the prompt is written by a human.

The generation is either text sampled from a language model, or the human-authored continuation to the prompt. Annotators, who do not know whether the generation came from a model or humans, assess this text. A paragraph length (80–145 tokens) is chosen to balance expressiveness with scope. For expressiveness, models must be given a sufficient number of tokens to express their capabilities lexically, syntactically, and semantically. One paragraph allows for significantly more variation than a single sentence. On the other hand, assessing multiple paragraphs is challenging, both as a crowdsourcing task itself, and because it broadens the kinds of errors to include larger narrative scope. We leave extensions of Scarecrow to longer narrative lengths for future work.

2 Span Labeling

Annotators select spans that contain problems in the generation. The spans are automatically snapped to word boundaries. We choose spans to balance specificity (i.e., vs. simply commenting on the text as a whole) with ease of use (vs. imposing a more structured annotation schema).

3 Span Selection

We instruct workers to select the smallest span—minimally a single word—that contains an issue. Sometimes this involves an entire phrase, sentence, or multiple sentences. We aim for specificity because during aggregation, it is possible to “back off” annotations to larger spans, but not the inverse.

Once they select a span, workers (1) label the error type, (2) choose a severity level, and (3) explain their reasoning behind the error. Workers use the annotation interface shown in Figure 5 to mark a span with these three steps. We describe each step in greater detail in the next three sections.

4 Error Types

Each selected span is labeled with exactly one error type. Multiple errors may be marked with partially or fully overlapping spans in the case that one text segment contains multiple problems.

We chose ten error types to balance three criteria: linguistic analysis, observed errors in generated text, and capabilities of everyday people with one to two hours of training.The complete training material is available for download. We developed the schema by starting with the first two criteria (linguistic analysis and observed errors), and refining it over several pilot annotation studies, with 30 crowd workers performing 750 total annotations of 60 paragraphs before beginning data collection.

We broadly group the errors into three categories: language errors, factual errors, and reader issues. Language errors are issues with internal and external structure of text: which ideas are expressed, and whether they are expressed coherently and consistently. Factual errors denote that the information presented is known to be incorrect. Reader issues, on the other hand, are cases where the text is too technical or obscure to assess its factuality. Hence, reader issues are not errors, per se, but regions where a reader would need assistance outside of the text itself for comprehension.

We present the ten error types in Table 1 (several pages back). Appendix A provides more details, examples, and explanations for all error types.

5 Severity

Errors naturally vary in how jarring they are to a reader. We define three error severity levels, and ask annotators to pick one for each error.

The severity levels are as follows. (1) Almost no impact on quality; just a small problem. (2) Understandable, but difficult; what’s written is still comprehensible, but there’s clearly an issue. (3) Very difficult to understand; the error almost completely ruins the text.

We provide examples of each severity in Appendix B.1. In this paper, we omit an analysis of the severity labels (except for an illustration in Figure 12), but include it in our data release for future work to explore.

6 Explanation

Finally, we ask annotators to explain their reasoning behind each error in natural language. We provide example explanations during training, but do not impose strict guidelines. This paper primarily focuses on quantitative error analysis, but we anticipate the error explanations may warrant future investigation.

7 Annotation Process

We use Amazon Mechanical Turk (AMT) for all data collection.

We first pay each worker 40totakeanextensivequalificationtask,whichbothtrainstheminthespancategorizationschemeandquizzestheirunderstanding.Wepassworkersiftheyscore40 to take an extensive qualification task, which both trains them in the span categorization scheme and quizzes their understanding. We pass workers if they score\geq$ 90 points out of 100 points (details in Appendix B.2).

Workers annotate each paragraph using a custom annotation interface (shown partially in Figure 5), for which we pay 3.50.Wecalculated3.50. We calculated3.50 per annotation by aiming to pay workers at least $15/hour. After several annotation rounds, we observed considerable variation in time per annotation,Median: 212s, mean: 265s, std. dev.: 199s. so this cost should not be necessarily seen as a requirement for Scarecrow annotations.

Data Collection

We collect 13k human annotations of 1.3k paragraphs using Scarecrow, resulting in over 41k spans.

We consider four model configurations to test recent state-of-the-art transformer-based Vaswani et al. (2017) models.

Radford et al. (2019) The 117M parameter variant of GPT-2, which is pretrained on WebText, without additional fine-tuning.

Radford et al. (2019) The 1.5B parameter variant of GPT-2, (WebText, no fine-tuning).

Zellers et al. (2019) The 1.5B parameter variant of Grover, a model with the same architecture and parameter count of GPT-2, trained on news articles and their metadata.

Brown et al. (2020) The 175B parameter variant of GPT-3, which is trained on a version of the Common Crawl web scrape with additional filtering and deduplicating.

In addition, we also use the actual human-written text from the data sources we draw from, which we denote as Human.

2 Decoding strategies

We consider three main hyperparameters when sampling from models: pp for top-p or nucleus sampling Holtzman et al. (2020), an alternative to top-k;We omit separate studies of top-k, due to results presented by Holtzman et al. (2020), and OpenAI’s removal of top-k from the GPT-3 API. t for the softmax temperature; and f.p. for frequency penalty. The frequency penalty scales a token’s likelihood based on how many times it was already generated by applying the following modification to the model’s output:

To compare models as consistently as possible, we set identical decoding strategies for our primary data collection. We refer to this as the “apples-to-apples” decoding setup throughout the paper:

However, we also wish to study the effects of these decoding strategies. We annotate generations from the strongest available model (currently, GPT-3) varying the following parameters:

For budget reasons, we only vary pp and tt independently—i.e., we set p=0.96p=0.96 when varying tt, and t=1.0t=1.0 when varying pp.

3 Prompt Selection

We use news articles as the sources of prompts for models to condition on for generation. Specifically, we use news articles found in the Common Crawl. We select the first sentence as the prompt.

Our use of news text is constrained by two factors. First GPT-3 is trained on the Common Crawl, from 2016 through 2019. We wish to avoid testing GPT-3 by generating from articles it saw during training, due to the possibility of copying Carlini et al. (2021). Second, news articles began heavily covering the COVID-19 pandemic beginning around February 2020. Though testing models’ capabilities to generate text about unseen events is a valuable line of study, the distribution shift caused by COVID-19 in news writing about all aspects of life is difficult to overstate.

As such, to make the comparison more amenable to models’ training data, we consider news articles from January 2020. We select articles where there is a known topic—such as Food or Sports—from the Common Crawl metadata, to allow for studying any effect of coarse-grained subject.

4 Generation

We generate between 80 and 145 tokensCounted by Stanza tokenization Qi et al. (2020), not byte-pair encoding (BPE) or whitespace-separated tokens. from each model as a continuation to the first sentence of the news article. We stop generating when we heuristically detect the first sentence boundary after 80 tokens. If the model does not end a sentence between 80 and 145 tokens, we sample again. For the Human setting, we use the remainder of the article, similarly stopping after the first sentence boundary after 80 tokens.

5 Annotation

Workers first complete training and qualification tasks. We provide more details in 4.7. From pilot studies, we discovered that each error, depending on its severity and clarity, has only a low to moderate chance of being identified by each worker. However, most worker-identified errors were truly problems. In other words, annotators labeled issues with high precision and low recall. To account for this, we have 10 workers annotate each paragraph. We examine the agreement and variability of annotations in Appendix C.

We provide detailed dataset statistics in Appendix D.

Error Prediction

A natural question is: using this data, can machines learn to detect and classify errors in machine generated text?

We frame this problem as a span classification task. Given a span from a generated text, the goal is to classify its error type or output “No Error” if there is none. Positive examples for each error class are taken from our data. We sample random spans that were not labeled with any error type as negative examples. To ensure a breadth of span lengths, we sample 3 negative spans for every length of error span in the generated text. We split the generated texts into train, development, and test sets using 1063 texts (28029 error spans), 100 texts (2538 spans) and 100 texts (2677 spans) respectively.

We use a standard span classification model inspired by Wadden et al. (2019). This model encodes every generated text using a pretrained language model (RoBERTa-large). Spans are represented with the final layer of this encoding. Following previous work, we concatenate the start and end tokens with a task-specific learned length embedding. The resulting vector is passed through a feedforward network which reduces its dimensionally to the number of error categories plus a “No Error” option. The resulting model has 357M trainable parameters. The model is trained to minimize the cross entropy of the correct span category. We train for 15 epochs using AdamW with a learning rate of 10−610^{-6}. We validate after each epoch and use the checkpoint with the lowest validation loss (epoch 8).

To evaluate the error prediction model, we use per-token precision, recall, and F1 score per error category. We classify every span up to length 30 in a generated text. We take as gold labels the aggregated human error spans collected in our data. In other words, models predict the combined spans of all 10 annotators. For comparison, we also report as Human the average metrics of one annotator versus the others (i.e., 1-vs-9).The difference in available references (10 for models, 9 for humans) mean this setup makes it easier for models to score higher in precision, and for humans to score higher in recall. Despite this, humans still achieve higher precision, and models still achieve higher recall.

Table 2 shows the error prediction capability of this model in terms of precision and recall. As we noted earlier, a single human annotator can be thought of as a high precision, low recall judge. These results bear out this claim. For all but one category, humans have higher precision annotations. However, the models trained on the aggregation of human labels can achieve considerably higher recall. For half of the error categories, this leads to higher model F1 scores than the human annotators.

We see that the model is successful at identifying information that human’s would have to manually verify ( [ Needs Google]), achieving nearly perfect recall with precision close to 0.6. The model can also identify [ Grammar and Usage], [ Incoherent], and [ Redundant] errors with higher recall than an individual human annotator, though at the cost of precision (sometimes in the .20s).

Related Work

Automated evaluation metrics such as BLEU Papineni et al. (2002), ROUGE Lin (2004), METEOR Banerjee and Lavie (2005), and BERTScore Zhang et al. (2019) compute a generation’s score based on a (set of) reference(s). Their use is well-established in tasks like machine translation and summarization, but they are less helpful in open-ended text generation, where there is a vast diversity of possible high-quality continuations.

Recent studies propose automated metrics for open-ended text generation evaluation such as: Perception Score Gu et al. (2021), which diffuses evaluation onto a multidimensional space and assigns a single score; UNION Guan and Huang (2020), which learns to distinguish human-written stories from negative samples by generating perturbations of human-written stories; and MAUVE Pillutla et al. (2021), which compares the distribution of machine-generated text to that of human language.

An alternate recent approach to assessing open-ended text generation was presented in TuringAdvice Zellers et al. (2021), where crowd workers assess machine-generated advice in response to Reddit posts. In their error analysis, Zellers et al. connect problems in generated text to core NLP tasks, such as [ Self-Contradiction] errors as instances of failed natural language inference Monz and de Rijke (2001), or [ Off-Prompt] errors as cases of failed reading comprehension Richardson et al. (2013). While past work has attempted to guide text generation using discriminative models trained for such tasks Holtzman et al. (2018), it remains an open challenge.

Comparative human evaluations of natural language generations ask annotators to rank system outputs relative to each other. Text is typically evaluated using a few global criteria, such as fluency and relevance, using discrete (e.g., 5-point) Sai et al. (2020) or continuous scales Novikova et al. (2018). Recent work even automates this approach, running a human evaluation alongside automatic metrics on leaderboard submissions Khashabi et al. (2021). In the RoFT system (Dugan et al., 2020), annotators attempt to detect the boundary between human- and machine-written text as a proxy for assessing quality. Table 3 summarizes the differences between these schemes and Scarecrow. See Celikyilmaz et al. (2021) for a recent survey of text generation evaluation techniques across both human and automatic metrics.

While these approaches may be helpful—sometimes Card et al. (2020)—at ranking systems, they do not give us insight into exactly which parts of a generation fall short, and why. One approach related to or annotation method is pursued by Wood et al. (2018), who develop a collaborative mobile app where users draw “graffiti” commentary on news articles. Scarecrow aims to assess model generations the way we would critique human-written text: by locating, coarsely categorizing, and explaining problems.

Conclusion

We present Scarecrow, a method for identifying and explaining issues in generated text. Along with the annotation framework, we present an analysis of the Scarecrow method applied to several large neural language models in an open-ended news generation task. We release our data and methodology to the community.

Acknowledgments

The authors thank members of xlab for their feedback on this work. This research is supported in part by NSF (IIS-1714566), DARPA MCS program through NIWC Pacific (N66001-19-2-4031), DARPA SemaFor program, and Allen Institute for AI.

References

Appendix A Scarecrow Annotation Schema

Here, we present in greater detail the Scarecrow annotation error types.All example annotations here are our own. Many are provided to annotators during training. A visual summary is shown in Figure 6.

While we annotate using this schema, the essence of our study is to embrace language users’ abilities to detect when something may be wrong with text. In other words, we do not wish for our span definitions to get in the way of humans describing problems with text. To this end, we encourage researchers to embrace label back off (to coarser categories), merging labels (based on empirical observations), and refining the annotation ontology over time. The central goal is to collect what people find wrong with text.

We define five categories of language errors, which concern the selection of ideas in a text and how they are expressed. These range from grammar and syntax problems to issues of semantics and pragmatics.

This category of errors includes missing words, extra words, and incorrect or out of order words.

example A PhD student from the University of Kent in the UK claims to have discovered a clever way to explain the positive [ emoticons] in cats. Explanation: The word should probably be “emotions.”

We also label [ Grammar and Usage]for inserted words or small phrases that could be deleted to resolve the issue:

A couple is facing criticism for their extravagant birthday party. The bewitching pair had first stripped down to fishnets [ and backward.] Explanation: This phrase can simply be deleted.

We avoid partitioning [ Grammar and Usage] errors into more detailed categories based on the observation that large language models produce fewer issues of syntax and diction (aside from [ Redundant] errors, described next). As such, we focus instead on semantic and pragmatic errors, captured by the upcoming error types.

A.1.2 [ Redundant]

While “redundant” can also include extra unnecessary information, we specifically use the [ Redundant] label to mark repetition. In identifying redundant text, our schema annotates both the [ antecedent] (first mention) and the [ redundant text] (when the repetition occurs). Sometimes the exact word or phrase will be repeated.

example Many merchants worry about the possibility of [ poor service] [ or service] for certain categories of customers.

Other times, generated text expresses the same idea repeatedly using different words.

example They then made decisions based on Kondo’s instructions, to the extent that they [ created de-cluttered spaces] [ and got rid of clutter and clutter-filled spaces].

A.1.3 [ Off-Prompt]

The prompt is a human-written sentence used as context from which the model generates a continuation. Models sometimes generate text that is unrelated to the prompt.

example Prompt: Dogs are the new kids. Generation: Statistics suggest that most Americans would be happier with dogs than children. [ In fact, four out of five don’t even visit the dentist annually, much less every six months.] Dog owners report much higher rates of happiness than non-dog owners.

Other times, the text may be related, but it contradicts what is stated in the prompt.

example Prompt: China sets new record for Economic Growth Generation: The Chinese economy [ fell 10% this month, the third such loss this year.]

A.1.4 [ Self-Contradiction]

When a model generates text that contradicts the prompt, that is labeled as [ Off-Prompt]. But when a model generates text that contradicts itself, that is labeled as [ Self-Contradiction]. We also mark the [ antecedent](original statement).

example McDonald’s is considering a design which will replace the [ cardboard packaging.] Mr Gore-Cotter said: “We recognise the concern around waste. We are now looking at a new design that minimises the [ plastic bag.”] Explanation: The idea of minimizing the plastic bag contradicts the stated goal of replacing cardboard packaging.

example Mall of America plans to [ lay off and furlough hundreds of its employees.] [ It has no plans to restrict the number of hours workers can work.] Explanation: Furloughed workers are explicitly restricted from working.

A.1.5 [ Incoherent]

Generated text is sometimes grammatical, not redundant, on prompt, and not contradictory, but still confusing. We provide the [ Incoherent]label for such sentences.

example Melody Mitsugi, 28, had never given her kids cheese toast before her husband [ drew a map of it on her toast.] Explanation: One can’t exactly draw a map of Cheese Toast, and one probably wouldn’t draw it on toast itself.

example Cats naturally show anxiety and fear by at times [ breaking apart different parts of the brain in an attempt to keep the others from escaping.] Explanation: It’s difficult to even imagine what is happening in this passage.

A.2 Factual Errors

We define three categories of factual errors, which encompass known incorrect statements.

Generated text will sometimes have issues with basic mathematical operations of known quantities (e.g., “half of ten apples is four”), problems converting fixed units (e.g., m to cm).

example One account, @Iain_Rowling1, had over 500,000 followers at one point, but in just four days they fell by [ around half - some 4,000.]

We also include problems converting currencies that are wildly implausible under modern assumptions (e.g., £1 = $18 US ).

example … compared with just over £1,000 [ ($18,868)] for previous versions of Samsung’s flagship phone.

A.2.2 [ Commonsense]

These errors mark spans that violate our everyday basic understanding of the world. Though it is challenging to precisely define commonsense knowledge Liu and Singh (2004), we include non-encyclopedic knowledge and basic reasoning.

The following example concerns broadly sensible numerical ranges.

example The picture is from high above the South Pole, where close to Astronauts live and work. Explanation: Even if we don’t know the exact number of astronauts in space, it is common knowledge that 100k is far too many.

The next example involves world knowledge, akin to scripts Schank and Abelson (1977).

example You can get the dress custom-made and stitched at your favorite [ spa.] Explanation: Spas don’t offer stitching.

The following example involves lexical entailment.

example The thinness of our bodies isn’t an answer to all common human health problems like [ obesity] or diabetes Explanation: While most of the statement is acceptable, it’s impossible to be “thin” and “obese” at the same time.

example Now in 2021, NASA is measuring California wildfire temperatures using an instrument on the International Space Station. This year’s record-shattering heat has had global repercussions in , forcing sea level rise on California and increasing the risk of deadly wildfires. Explanation: Events in 2021 can’t affect events in 2017.

A.2.3 [ Encyclopedic]

These errors are ones that we know are factually wrong, and that we could look up in, say, Wikipedia.

example [ Japanese Prime Minister Justin Trudeau] said he will be halting all imports and exports until the current situation can be contained. Explanation: Justin Trudeau is the Prime Minister of Canada, not Japan.

The distinction between [ Encyclopedic] errors, and the upcoming [ Technical Jargon] and [ Needs Google] issues, depends on the reader’s knowledge.

example The gas contains something known as [ phyto-romatic acid, a common chemical element in the periodic table.] Explanation: Acids aren’t elements.

A.3 Reader Issues

We define two categories of reader issues. These are words or statements a reader cannot verify without using an external resource.

Sometimes generated text includes specific words from a field that requires expertise to understand.

example In Chile, an 800-megawatt [ photovoltaic] plant was built for a record low cost of $129 per megawatt-hour last year.

Which words are jargon depends on the reader’s particular expertise. This means [ Technical Jargon] spans are more accurately thought of as potential issues rather than known errors.

example He uses a spirit [ mash] made from white corn and malted barley and a [ neutral grain], which he describes as a "whiskey grain.”

A.3.2 [ Needs Google]

Many facts—especially those involving specific people, events, dates, or numbers—could be categorized as encyclopedic knowledge. However, whether the fact is accurate may require additional verification by the everyday reader. To make this distinction between known encyclopedic knowledge and trivia, we introduce this label to denote that a reader would need to search online to verify whether it is true.

We instruct annotators to not look up facts marked with the [ Needs Google] span. We do this to keep the focus of the task on classification, rather than factuality detection. As a result, [ Needs Google] spans mark statements that would need to be verified, rather than known errors.

example It was promoted by [ Dr. Michael Fanning, the Executive Director of the Foundation for Mental Health Awareness, Inc.] Explanation: A reader would likely need to look up whether there is a Dr. Fanning who holds this position.

example … an [ 800-megawatt photovoltaic plant] was built for a [ record low cost of 129permegawatt−hour]lastyear.Explanation:Inadditiontopotential[TechnicalJargon]spans,thereareatleasttwo[NeedsGoogle]spans:1.whethersuchaplantcanberoughly800−megawatt,2.whether129 per megawatt-hour] last year. Explanation: In addition to potential [ Technical Jargon] spans, there are at least two [ Needs Google] spans: 1. whether such a plant can be roughly 800-megawatt, 2. whether129/megawatt-hour is a sensible cost measure, and the value is reasonable.

To illustrate the annotation methodology and schema in practice, we present four complete example annotations in Figure 7. This figure also illustrates how much variation we see across models.

Appendix B Annotation Details

We provide here examples for each of the three error severity levels, which we also give to annotators during training.

example Paul Campbell-Hughes, from the University of Aberdeen, explains how [ she] managed to locate colonies of honey bees in Kent. Severity: 1. Since Paul is usually a male name, the model should have used “he.” But this error is pretty minor.

example Paul Campbell-Smith, a PhD student from the University of Kent in the UK, claims to have discovered a clever way to explain the positive [ emoticons] in cats. Severity: 2. The word should probably be “emotions.” We can guess what was being said, but it’s definitely wrong.

example Prompt: Whether you’re on Facebook, Instagram, Snapchat or TikTok, many people make huge efforts to curate the best version of themselves online. Generation: [ This year we’ve got something for you: a Love Match Custom Size Poster featuring Mather, Phoenix, Kashun and all her friends, divided among six different covers, creating a beautiful custom size poster for your own personal high school reunion.] Severity: 3. Even ignoring the end of the generation (a poster for a personal high school reunion?), this whole generation is way off the prompt and does not make sense.

B.2 Grading Details

In the training material, there are 10 annotation exercises, 10 multiple choice questions, and 1 real task question to test workers’ understanding.

After going through each error type, there is an annotation exercise. Workers are asked to mark the span with that particular error in a short text. Each exercise is worth 5 points.

After going through all language errors, and going through all factual errors and reader issues, there is a language error label quiz and a reader and factual error label quiz respectively. Each label quiz consists of 5 multiple choice questions, where workers are asked to choose the error type of a marked span in a short text. Each multiple choice question is worth 3 points.

At the end of the whole training material, workers are asked to apply what they learn in an actual task where they annotate a given paragraph with full tool like ones shown in Figure 7. This question is worth 20 points. We mark 7 error spans as the solution. As long as they can mark 5 of 7 error spans, they get a full 20 points. Otherwise, 4 points will be deducted for each missing error span.

In total, there are 100 points. We pass workers if they score ≥\geq 90 points, and then they are provided with the solution to review.

Appendix C Data Quality

Identifying and classifying errors in potentially noisy machine-generated text is a challenging task. How consistent are the annotations collected from crowd workers? In this section, we examine the agreement and variability of the collected annotations.

At a high level, we observe either acceptable or high inter-annotator agreement across error categories. For rare error types such as [ Bad Math], high agreement stems from the prevalence of spans with no error. For such categories, we recommend treating each annotator as a high precision, low recall judge, and considering the information from their aggregate annotations. Figure 8 gives an example of the perspective gained by viewing all 10 annotations of a single generation.

Table 4 shows token-level inter-annotator agreement statistics aggregated over all collected data. Since a single annotator can label a single span with multiple errors, we break the agreement statistics down by error category. We report Krippendorff’s α\alpha coefficient, a chance-corrected measure of agreement for multiple annotators (Krippendorff, 2018). Due to computational constraints, we calculate this coefficient per generation and report the average across the dataset. The agreement shown here is high for most categories (>>0.8) and acceptable (>>0.6) for all error types.

The Krippendorff measure may be deceptively high for some error types such as [ Bad Math], where 99% of tokens are not annotated with this error. The Two Agree measure in Table 4 gives a different characterization of this data. Two Agree for a given error label is the percentage of tokens labeled by at least one annotator that were also labeled by one or more additional annotators. This metric allows us to see where annotators agree that particular errors exist while ignoring the majority of tokens (for most error categories) which annotators agree are not errors. Two Agree shows significantly lower rates for sparse errors with high Krippendorff scores, such as [ Encyclopedic]. However, it reveals stronger agreement among [ Incoherent] and [ Off-Prompt] errors than might be expected given the Krippendorff coefficient.

A limitation for both metrics is the use of token-based overlap.

One issue we face is high variance of annotations. To determine the impact of this variance for lower-data settings, we perform a bootstrap analysis using largest subset of our data (GPT-3, top-p=0.96p=0.96, t=1t=1, f.p.=0=0, for which we have annotations of 200+ generations). We choose 50 generations (roughly 500 annotations) and calculate the error statistics therein. We repeat this process 1000 times and report the mean, standard deviation, and coefficient of variation in Table 5. We also calculate the coefficient of variation for different numbers of samples, shown in Figure 9. We see that as the number of samples increases, the coefficient of variation decreases as expected, though less precipitously after 30 examples. These results show that with as few as 50 documents, the Scarecrow error analysis should yield relatively robust results. However, this varies by error type: rare errors like [ Bad Math] and [ Encyclopedic] show greater variance. Here, again we repeat our recommendation to treat annotations for these categories in aggregate. These results motivate our collection of at least 500 annotations per condition studied.

Appendix D Dataset Statistics

We list the data collection quantities in Table 6, and plot visualizations of three aspects: prompt topic and annotated span proportions are shown in Figure 10, and average span lengths are shown in Figure 11.

Appendix E Detailed Analysis

In this section we perform a detailed analysis of the trends of individual error types and decoding configurations.

To begin, we consider apples-to-apples model decoding configurations. To expand on these results, originally presented in Figure 2, we also present two additional ways of counting error spans, which we show in Figure 12. While our method for counting errors throughout the paper takes into account the number of tokens covered in each span (span coverage), we also show plots for scaling each span by its severity level (span coverage ×\times severity), and by ignoring both severity and token length (simply span counts). These changes in measurement further illuminate model error characters, which we discuss in the upcoming sections (refer to Figure 12).

Under initial analysis of span coverage, [ Off-Prompt] errors show a model plateau at GPT-3. Measuring span counts offers barely perceptible improvement, indicating that scaling language models over more in-domain training does not guarantee topicality.

This observation is consistent with growing work on prompt programming as a new technique for attempting to steer large pretrained models to complete the desired task Branwen (2020); Gao et al. (2020); Reynolds and McDonell (2021). In practice, we observe that while GPT-3 will sometimes continue a prompt by writing an article, other times, it may elaborate on the prompt itself:

prompt Do you prefer the idea of being outdoors in the fresh air to being stuck inside with phones ringing and messages pinging? GPT-3 Can you leave work at work? Are you flexible enough to [ cover holidays or take on additional responsibilities? Can you prioritize tasks? If your boss comes to you on Tuesday to confirm the new social media strategy, are you able to pick up the ball and get the messaging hammered out by Thursday?] …

Of course, this generation is not literally [ Off-Prompt], but it is out of place when other generations are continuations of the prompt, rather than further elaborations of it.

While avoiding [ Off-Prompt] errors for language models is worth exploring with prompt programming and other avenues, an investigation of these techniques is outside the scope of this work.

Finally, we note that [ Off-Prompt] spans are the most prevalent error (not reader issue) marked for human-authored text. We suggest that a higher rate of false positives for this error type, coupled with its prevalence in model-generated text, makes further refinement of this error a compelling avenue for further study.

E.2 [ Self-Contradiction]

While changing from span coverage to span counts alters the relative order of GPT-2 XL and Grover (though still within confidence bounds), the puzzling question is why GPT-2 Small performs better than most (or all) other models. Why would the smallest model produce the fewest [ Self-Contradiction] errors?

We posit the reason is that GPT-2 generations are so [ Incoherent] and [ Off-Prompt] that there is little opportunity for relevant, comprehensible points to be made and then reversed. For example, see the GPT-2 Small annotated generation in the top left of Figure 7. The entire text is covered by [ Off-Prompt] and [ Incoherent] errors.The high double-error coverage reveals another consideration: to what depth (i.e., number of overlapping spans) will annotators mark? By the design of our framework, [ Incoherent] errors serve as a fall-back, but without it, we might imagine poor generations splatter-painted by other error types. If we look at GPT-2 Small’s error distribution in Figure 3, we see most of its added density comes from significantly more [ Off-Prompt] and [ Incoherent] tokens.

E.3 [ Redundant]

The different counting methods shown in Figure 12 reveal a change in the results for [ Redundant] errors. Rather than repetition simply increasing as models grow larger, we observe that GPT-3 repeats in a similar number of cases (lower span counts), but for more tokens (higher span coverage). This matches the qualitative observation that GPT-3 produces larger topically repetitive blocks, rather than simple word or phrase repetitions generated by GPT-2-sized models:

GPT-2 Small … owners have started growing their own breeds and dogs are [ starting to start] so there’s really … GPT-3 The focus of your thoughts should be on the task at hand, [ not on your productivity. You shouldn’t be thinking about how you can be more productive. You should be thinking about how you can be productive right now.] …

Such repetitions can be more difficult to clearly isolate, because even slight wording changes produce variations in tone and connotation. Rather than being identical semantically, we observe GPT-3 will seem stuck on a particular topic, elaborating on and rephrasing similar ideas more times than a human writer (hopefully) would.

E.4 Reader Issues

We observe the highest number of [ Needs Google] and [ Technical Jargon] issues in human-authored text.

[ Needs Google] issues broadly represent any specific claim that could be fact-checked. In our domain (news articles), these are primarily whether an event happened on a particular day, whether a person holds a role, or whether a mechanism works as described (e.g., chemical or technical). As seen in Figure 13 (which shows GPT-3’s span distribution), [ Needs Google] issues happen roughly equally for all topics. We believe this trend is due to the news article domain, which is prone to a high density of specific information. As such, for other domains, this trend may be less prevalent, more difficult to label (e.g., subtle claims assumed to be true in long running text), or both.

We observe that [ Technical Jargon] issues are influenced by topic (Figure 13, bottom), occurring significantly more frequently in Business, Health, Science, and Technology topics than in others. This trend displays a clear topic-dependence even within a single broader domain (news). These results indicate that both reader issues are characteristics of natural text. Of course, one might wish to measure or minimize potential reader issues for a particular application—for example, claim verification, or controlling for reading level.

E.5 Decoding Hyperparameters

We discuss the effects of the decoding hyperparameters we consider—top-pp, temperature, and frequency penalty—on generation quality. For the sake of annotation cost, we only vary these parameters for the strongest model available, GPT-3.

First, we show the effect of varying top-pp and temperature alone (i.e., with no frequency penalty) on different error types. Figure 14 shows the effect on two salient spans: [ Off-Prompt] and [ Redundant]. (We omit others for space.) We observe that annotators naturally label errors the way we would intuitively expect the model to produce them, given the hyperparameter changes. The bottom-right corner of each subplot, where t=1t=1 and p=0.96p=0.96, is the configuration with the highest amount of randomness from sampling. As we move away from that corner—either left by lowering temperature, or up by lowering top-pp—we lower the amount of randomness. We observe a positive correlation with randomness and [ Off-Prompt] errors, and an inverse correlation with [ Redundant] errors. In other words, sampling from a larger set of words makes the model more prone to changing topics, but less likely to repeat itself, and vice versa.

After confirming these intuitive measures, we turn our attention to Figure 15, which investigates the overall error spans for GPT-3 both without (left) and with (right) the frequency penalty. (Note that unlike Figure 14, both heatmaps in Figure 15 have the same color scale.) We observe that introducing the frequency penalty lowers error rates for every value of temperature and top-pp that we try. Furthermore, it appears to reverse the trend seen without a frequency penalty: that sampling from a larger set of words produces fewer errors.

The overall results for all decoding configurations were shown previously in Figure 4. In the next section, we focus on the GPT-3 decoding configuration that produced the fewest number of errors, and compare it to human authored text.

E.6 Best GPT-3 vs. Humans

The best GPT-3 configuration shown in Figure 4—argmax sampling with frequency penalty = 1—appears to match error rates seen in human text. Is the text generated by this model truly as error-free as news articles?

We first look at the error composition of both sets of annotations. To get a clear picture of the potential problems, we plot only error spans (ignoring reader issues), and we omit length scaling, instead plotting span counts. This breakdown is shown in the left plot of Figure 16. The error compositions are similar, the largest differences being more [ Redundant] errors for GPT-3, and more [ Grammar and Usage] errors for human-authored text.

Next, we perform a manual analysis of 160 errors, sampling 10 at random from each of the 8 error types for each model (GPT-3 and human-authored text). We show the results in the center plot of Figure 16. We notice that a greater portion of errors in human-authored text were due to artifacts present in the text-only format of the Common Crawl. For example, links to other articles or advertisements sometimes appear in the middle of an article’s text. While annotators were quick to mark these spans, they reflect errors in formatting, not in writing. We partition these errors separately and exclude them from the subsequent calculations.GPT-3’s generations also sometimes exhibited what appeared to be formatting errors due to training on web-scraped text, though more rarely. For example, some generations contained Which? after vague noun phrases, which appear to be learned from Wikipedia, where under-specified information is tagged by an editor with this word. For fairness, we removed these errors from GPT-3’s tally as well, though they were few enough we do not plot them separately.

Finally, we scale each error type’s prevalence for each model (i.e., the left plot of Figure 16) by the portion of errors that we estimate to be legitimate based on our manual annotation (i.e., Figure 16, center) to produce the right plot of Figure 16. After taking into account each error type’s frequency, we estimate that 48% of GPT-3’s worker-annotated errors overall are legitimate, compared to 9% for human-written articles.

This analysis suggests two findings. First, human-authored news paragraphs contain many times fewer issues than text authored by GPT-3 using the best decoding configuration we tested. Second, the noise of error annotations may be as high as 90% when assessing high-quality text. Though it would require further manual annotation to verify, we conjecture that the trend of GPT-3’s error spans being more reliable (only 50% noise) would continue, and that text generated by GPT-2 would contain even fewer false positives. We note that such rates are not fixed—after all, the manual annotations were done by one of the authors simply by reading carefully—but that more realistic text may require correspondingly more effort by human annotators.

E.7 Topics

As noted in §5.3, we collect data using prompts drawn primarily from 12–14 news topics. For conciseness, we show results only for GPT-3, and only for the standard apples-to-apples decoding configuration.

Figure 17 plots, based on the prompt topics, the average portion of the generation that is covered by error spans. While there is no significant difference between most topics, the results do indicate that generating text in more technical domains leads to higher span counts.

Figure 13 shows individual span prevalence by topic. The top heatmap normalizes each topic (column) independently. [ Needs Google] issues and [ Off-Prompt] errors dominate the error types, with a few exceptions: for History, and Nature articles, [ Redundant] trumps [ Off-Prompt] as a source of errors.

For the bottom, if we instead normalize by error label (row), we can observe which topics are more prone to certain error types than others. For example, we can see [ Bad Math] errors are most common in Business and Health generations; Entertainment causes the most [ Self-Contradiction] errors; and [ Technical Jargon] issues appears more frequently in articles about Business, Technology, or Health.

E.8 Error explanations

Figure 18 displays word clouds for common unigrams and bigrams found in the error explanations for each error type, and Figure 19 shows the average explanation lengths for each error type. For [ Technical Jargon], [ Redundant], and [ Needs Google] error types, the prominent words do not provide much illumination and they have short average explanation length, indicating that the explanations are straightforward affirmations of the category (“I think this is financial jargon,” “The information is repeated,” or “I would need Google to check this.”). But for categories like [ Encyclopedic] and [ Bad Math], we observe some coarse trends: “year” is prevalent in both, “movie” appears in [ Encyclopedic], and “million” is present in [ Bad Math], which suggests that the explanations are more likely from outside knowledge and needs some calculation (“The iPhone uses a lightening connector not a L-shaped connector,” or “5000 feet is 1524 meters.”)

Figure 20 presents a few representative explanations for four error types, taking particular note of their explanation lengths (Figure 19). Both [ Self-Contradiction] and [ Redundant] errors have antecedents, but their explanations are markedly different. Explanations for [ Self-Contradiction] contain more information describing the particular semantics that is reversed, which are less obvious at first glance than other errors. On the other hand, [ Redundant] errors are more straightforward to spot, often involving simple lexical overlap, and so don’t require elaboration.

Explanations for [ Commonsense] contain the true commonsense knowledge that the text violates, which may take several words to explain. But an explanation for a [ Grammar and Usage] error simply corrects the error; as these errors are easier to fix, the explanation lengths are often short.

Appendix F Future Work

We outline several further directions of study centering around the Scarecrow annotation framework, considering both natural implications and broader steps.

We observed that for GPT-3, a frequency penalty value of 1 with argmax sampling produced fewer error spans than any other configuration (Fig. 4). We have not tried varying the frequency penalty to values between 0 and 1, or adding any presence penalty (§5.2), both of which then allow for fresh explorations of top-pp and temperature.

How good can (a finetuned) GPT-2 get? We saw decoding parameters considerably impacted GPT-3’s performance, moving it from edging out Grover to error rates close to humans (Fig. 4). Could such decoding changes have a similar effect on a GPT-2-sized model? Or might a smaller model favor different decoding hyperparameteres?

We observed good annotator agreement given the complexity of the task, but the odds that two annotators agree exactly on each span’s type and boundaries remains only moderate (§C). We did not try backing-off (a) error types into coarser categories (e.g., language, factual, reader issue) or even to binary presence; (b) span boundaries into phrase or sentence-level annotations. Applying a type of back-off could also allow clustering methods to discover different error ontologies.

While we present baseline results for automatic span error detection (§6), we anticipate that significant progress is still available in this new task.

F.2 Scarecrow Studies: Complex

In the current work, we largely treat annotators independently, with the exception of measuring their overlap to study agreement (§C) or taking their union to train prediction model (§6). However, we might consider other ways of viewing the 10 annotations for each generation together. For example, we might consider the aggregate decision of whether a token is labeled with any span a measure of how noticeable or jarring an error is. This measure may be related to error severity, but may be distinct from it.

One might also consider formal methods for computing annotation alignments. The Gamma measure, proposed by Mathet et al. (2015), satisfies the long list of criteria needed to align and measure Scarecrow annotations: spans of multiple types, with gaps, full and partial span overlap, more than three annotators, and the potential to merge or split annotations (which we have not addressed in this paper). While we performed experiments with this measure, we experienced difficulties producing intuitive alignments with the authors’ software, which disallows configuring parameters of the mixed-integer programming problem.The mixed-integer programming approach is also computationally intensive; e.g., memory alone prevented us from computing alignments for pilot studies with twenty annotators, even on a machine with 500GB of RAM. Emerging concurrent work Titeux and Riad (2021) offers a reimplementation of this measure that exposes additional parameters, which may be a promising avenue. However, it is possible that aligning annotations is a challenging task on its own that might require use of the explanations.

Related to the previous point about error alignment, one might study whether model size affects span agreement. Anecdotally, errors from larger models like GPT-3—even of the same type, like [ Commonsense] errors—are more difficult to describe without careful consideration, and may also be more difficult to identify.

Our quantitative studies of [ Redundant] errors (e.g., Figs. 14 and 12) point to semantic repetition as the major issue that emerges as models are scaled. Though this effect may be mitigated by changes to the decoding algorithm (like the frequency penalty), we still observe that models have difficulty striking a balance of repetition. With excessive paraphrasing, generated text seems stuck on an idea. But equally, if a generation moves too quickly between ideas without linking them together or to an overall theme, the text lacks coherence. We posit that the issue of [ Redundant] text emerges as the shadow of encompassing issues of narrative structure and discourse.

F.3 Broadening Scarecrow

This paper focuses on open-ended generation, but a natural extension of this method would be to assessing constrained generation tasks, such as machine translation.

Especially if considering a novel task setting, new error types may prove useful. For example, in constrained generation, one might consider an Adequacy error, which—as in machine translation—would indicate that the meaning of a span diverges from what is expected given the generation constraints. Furthermore, one might need to introduce annotations on the provided (not generated) text to account for desired semantic components that are missing from the generated text. Or, perhaps for a dialog setting, one might introduce a Generic label, which would indicate that a portion of the generation is otherwise coherent and correct, but offers a lack of new information.Such generic language may be seen as violating Grice’s Maxims Grice (1975), for example, by providing a dearth of information quantity, or by flouting improper manner by lacking brevity.

Other work has considered the evaluation of natural language generations at-scale, looking at distributional properties of the text Caccia et al. (2020); Pillutla et al. (2021). We suggest that these views are complementary to instance-based, human evaluation proposed here, and combining the approaches could lead towards a more holistic view of generative evaluation. For example, while all [ Self-Contradiction] errors right now are within-document, one could similarly identify cross-document contradiction errors, where a model is inconsistent at a more global scale.

F.4 Applications

One potential application of the Scarecrow data could be using the [ Needs Google] spans as a dataset of its own. In addition to training models to identify spans that require verification, one could go a step further and consider evidence retrieval for each span, and even propose a classification task.Minimally, [ Needs Google] spans from human-authored reputable news text should (hopefully) all be factually correct.

One errors can be detected, can they be fixed? The difficulty and scope of fixing Scarecrow-identified errors may depend on the error type, as error fixes may have cascading effects in the rest of the document.