Measuring Attribution in Natural Language Generation Models

Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins, Dipanjan Das, Slav Petrov, Gaurav Singh Tomar, Iulia Turc, David Reitter

Introduction

Large, pretrained neural models have advanced Natural Language Generation (NLG) performance across a variety of use cases, including text summarization, translation, and dialogue. Yet, generative neural models are known to hallucinate often, lacking faithfulness to underlying sources, for example in summarization or in grounded dialogue systems. Accurate evaluation with respect to these issues is important.

In this paper, we develop a framework for the evaluation of attribution, by which we mean the accurate use of source documents to support generated text. Attribution is closely related to issues of hallucination and faithfulness (see §2 for discussion). As a key motivating example, consider a dialog with a system that generates responses to a user’s sequence of questions:

If such a system, in addition to generating responses, could attribute its statements to source documents, that is, provide sufficient and concise evidence for its claims, system designers and users alike could more readily ascertain the extent to which the information it provides is supported by underlying sources. Prior work in NLG spanning diverse use cases such as summarization, dialogue response generation and data-to-text generation have investigated issues of faithfulness and “hallucination”, but have not provided a uniform and formally expressed framework to measure these errors. We discuss the relationship of our work to related work in §2.

In §3, we introduce our evaluation framework, Attributable to Identified Sources (AIS), that can be used to assess whether statements in natural language made by a system are derivable from a given underlying source. The definition of AIS (see §3.3.1) formalizes the meaning of a sentence ss in context using the notion of explicatures Carston (1988); Wilson and Sperber (2004)For example, in the above dialogue the explicature of “he was 25 years old” is “George Harrison was 25 years old when ‘Wonderwall Music’ was released”: the latter explicature is evaluated for attribution. Note that this use of explicatures is closely related to prior work on decontextualization Choi et al. (2021), see §2 for more discussion., and defines attribution to some background information source PP in terms of an intuitive test, asking whether “According to P, s”. It also accommodates system outputs whose meaning is uninterpretable. AIS can be used as a pre-condition or in tandem with other metrics or evaluation frameworks to assess overall quality. For example, characteristics of the underlying source (such as “source quality”), the fluency of the generated text, and so forth, can be measured using complementary metrics that are out of scope in this work.

We propose specific instantiations of AIS for three NLG tasks (§4): response generation in a conversational QA setting (as in the example above; responses must be attributable to a provided answer document), text summarization (where the summary must be attributable to the source article), and description generation from structured tables, or table-to-text (where the description must be attributable to the source table and associated metadata). Each domain involves a number of challenges: for example, in dialogue systems a key challenge is that the meaning of system responses is highly contextually dependent.

Next, we establish the feasibility of AIS evaluations via an empirical study through conducting human evaluation experiments. We train annotators to evaluate output text from multiple models per task using task-specific instantiations of AIS. We show that in our human evaluation studies, it is possible to achieve a moderate to high degree of inter-annotator agreement (see §4 for more details). We’re also able to observe differences in model outputs’ AIS scores, following generally expected trends. As part of this work, we release the detailed guidelines for human evaluation. We believe that AIS as a framework would be essential for the evaluation of system-generated utterances across NLG tasks.

Background

Hallucations in NLG. As alluded to in §1, past work has identified the issue of hallucination in neural generation models. Wiseman, Shieber, and Rush (2017) presented challenges in data-to-text generation where neural models generate hallucinated content not supported by source data; they proposed an automatic information extraction-based metric to evaluate generated text for that particular scenario and conducted a small human evaluation study examining whether summaries are supported by source data. More recently, Parikh et al. (2020) presented a larger human evaluation study in the context of a data-to-text generation dataset entitled ToTTo, and measured hallucinations in terms of faithfulness with respect to a source data table.

Hallucination has been a salient subject of investigation in text summarization. Maynez et al. (2020) presented an extensive characterization of hallucinations, discussed behavior of models that generate content that are present in larger corpora beyond a given source and conducted a significant human study. Additional automatic QA-based methods for detecting hallucinations have been proposed by Wang, Cho, and Lewis (2020), Nan et al. (2021), among others. One of the most relevant papers, Durmus, He, and Diab (2020), involved both a human evaluation and the introduction of an automatic question–answer based evaluation method. Their human evaluation of summary sentences is similar to our two-stage annotation pipeline where they evaluate sentences in two steps — first for whether it is understandable, and second, if so, for faithfulness to the underlying source (their instructions to annotators are: “If the information conveyed by the sentence is not expressed in the source, select ‘unfaithful’.”).

In the case of response generation for dialogue, especially in scenarios that involve the system responding about the real world, research has focused on measuring the responses’ consistency to prior conversational history or their groundedness to some external evidence, that we deem to be very close to the topic of hallucination. These have been measured via dialogue-specific natural language inference methods, often via human studies and data creation Welleck et al. (2019); Mehri and Eskenazi (2020); Gupta et al. (2021); Honovich et al. (2021); Dziri et al. (2021); Santhanam et al. (2021).

Despite a significant amount of work pertaining to hallucination spanning multiple NLG problems, there is no unified approach to evaluate whether system generated statements are supported by underlying source documents. Human evaluation studies are varied from paper to paper and detailed, reproducible annotation instructions are unavailable Belz, Mille, and Howcroft (2020). Likewise, the use of terminology for describing and defining evaluation criteria also lacks consistency and further complicates reproducibility Howcroft et al. (2020). General-purpose benchmarking across these tasks have gained traction Gehrmann et al. (2021), but there has not been a standardized treatment of the attribution problem. Our paper attempts to address this gap by explicitly formalizing the evaluation of attribution as a replicable and extendable conceptual framework. As part of our definition of attribution, we outline a more formal background for “information conveyed by the text” — in particular through the use of explicatures (see Figure 1 for examples). Lastly, we demonstrate that AIS can be generalized across multiple NLG tasks in which context, source documents, and generated text can take different forms.

Fact Verification. A related field of study has dealt with the topic of fact or claim verification (Thorne et al., 2018; Thorne and Vlachos, 2018; Thorne et al., 2021, inter alia). Work in this area has framed the task as retrieving supporting evidence given a claim, and optionally classifying semantic relationships between the claim and the evidence text. Modeling approaches have overlapped with recent literature examining natural language inference Nie, Chen, and Bansal (2019). Thorne et al. (2021) have examined several human annotation tasks for the above family of problems; however, there are several key differences with this work. First, we evaluate the quality of a system generated utterance with respect to given evidence source (a fundamentally different end goal); we utilize the notion of explicatures in defining attribution; finally, we avoid absolute judgments regarding “factuality” of utterances. As mentioned in §1, rather than making factuality judgments, we deem that complementary evaluation methods such as “source quality” in tandem with AIS would be required to evaluate the factuality of utterances. As a corollary, we assume the source is a reference, and that an actual system may select sources for their trustworthiness.

Decontextualization. Choi et al. (2021) introduce the task of decontextualization, that is, the problem of taking a sentence in context and rewriting it in a way that it’s meaning is preserved, and it can be interpreted out of context. This is directly related to the idea of explicatures, which are also used in the current paper.

A Formal Definition of Attributable to Identified Sources

This section gives a formal definition of AIS, attempting to give a clear and precise definition of attribution. We first give a definition of AIS for a simple case, where the utterance from a system is a standalone proposition. In spite of the simplicity of this setting, it is highly informative, and forms the basis for the full definition of AIS. We then describe how this definition extends to a much larger set of system utterances, in particular giving a treatment of interpretabilityWe acknowledge that the term “interpretability” has come to signify “model interpretability” in the NLP and ML community (as established in Harrington et al. (1985), Ribeiro, Singh, and Guestrin (2016)). The term in our use represents how interpretable system output is for a human annotator. The choice of terminology is intended to be more conceptually transparent when used by annotators: unlike other terms like “meaningful”/“nonsensical” Durmus, He, and Diab (2020), or “sensibleness” Adiwardana et al. (2020), “interpretability” more readily alludes to the significance of the propositions in system generated output in relationship to context. Finally, the annotators are typically not familiar with the “model interpretability” usage of the term., and contextual effects. A key idea in our model of meaning in context is the notion of explicatures Carston (1988); Wilson and Sperber (2004); Choi et al. (2021). In a final subsection, we describe how key aspects of the AIS definition naturally lend it to operationalization, while also pointing out how certain idealizations (e.g., the notion of a “generic speaker”) must be relaxed to accommodate the practical realities of implementation.

We now give a definition of AIS for a simple but important case, where the text in question is a standalone proposition. We in general assume a setting where AIS is to be determined for a string whose meaning is ascertained relative to a context. In the following treatment we assume that time is the only non-linguistic aspect of context relevant to determining textual meaning, modeling a setting where two generic speakers communicate over a text-based channel, with no additional prior information about each other.Extensions of AIS to more complex settings may require a more elaborate notion of non-linguistic context.

We define standalone propositions as follows:

A standalone proposition is a declarative sentence that is interpretable once a time tt has been specified.

To illustrate the definition of standalone propositions, consider the following examples:

George Harrison was 25 years old when his album ‘Wonderwall Music’ was released.

All four examples are declarative sentences. S1 is a standalone proposition. S4 is a standalone proposition, as it is interpretable once the time tt is specified. S2 is however not a standalone proposition, as it cannot be interpreted without additional contextual information: It is unclear what “He” refers to. More subtly, S3 is also not a standalone proposition, because it lacks details of historical context.

The definition of AIS for standalone propositions is as follows:

A pair (s,t)(s,t) consisting of a standalone proposition ss and a time tt is Attributable to Identified Sources (AIS) iff the following conditions hold:

The system provides a set of parts PP of some underlying corpus KK, along with ss.

A pair (s,t)(s,t) is attributable to a set of parts PP of some underlying corpus KK iff: A generic hearer will, with a chosen level of confidence, affirm the following statement: “According to PP, ss”, where ss is interpreted relative to time tt.

Here, the corpus KK could be a set of web pages, and the parts PP could be pointers to paragraphs or sentences within KK; or the corpus KK could be a knowledge graph, with PP as parts of the underlying knowledge graph; other examples are no doubt possible.

As an example, consider standalone proposition S1 given above, assume that the corpus KK is all of Wikipedia, t0t_{0} is the present time (specifically, noon on December 21st 2021), and assume that the set PP consists of a single paragraph from Wikipedia, as follows:

George Harrison (25 February 1943 — 29 November 2001) was an English musician, singer–songwriter, and music and film producer who achieved international fame as the lead guitarist of the Beatles. His debut solo album was ‘Wonderwall Music’, released in November 1968.

Under this definition, it would be correct for a hearer to judge “(S1,t0)(S1,t_{0}) is attributable to P1”, because the “according to” test in the AIS definition holds. That is, it is reasonable to say according to P1, S1” where S1 is interpreted at time t0t_{0}: “according to P1, George Harrison was 25 years old when his album ‘Wonderwall Music’ was released.”

Note that in some cases the system may provide multiple parts. The standalone proposition SS may also be justified by certain forms of multi-hop reasoning (e.g., arithmetic processes) over that set of parts. The above example requires reasoning about dates and age.

2 Extending AIS: Attribution of Sentences in Context

We now extend the previous definition of AIS to cover sentences that that go beyond standalone propositions. To do so, we will need to consider multi-sentence cases, and cases with non-empty linguistic contexts. We will also cover cases that are uninterpretable.

We first define the notion of “utterance”:

An utterance is a sequence of one or more sentences produced by a system or user, where a sentence may be a declarative, a question, a command, an exclamation, or a fragment. The iith system utterance is si=si,1…si,∣si∣s_{i}=s_{i,1}\ldots s_{i,|s_{i}|}, where si,js_{i,j} is the jjth sentence within system utterance sis_{i}, and similarly the iith user utterance is ui=ui,1…ui,∣ui∣u_{i}=u_{i,1}\ldots u_{i,|u_{i}|}.

To briefly illustrate our approach to non-empty linguistic contexts, consider the following interaction between a user and system (originally given in the introduction; repeated here for convenience):

The system utterance s2=\emhe was 25 years olds_{2}=\hbox{\em he was 25 years old} is clearly not a standalone proposition. As such, it cannot be evaluated for AIS given our previous definition. However, given the previous context in the interaction, intuitively the meaning of s2s_{2} is something similar to the standalone proposition “George Harrison was 25 years old when his album “Wonderwall Music” was released”. This latter “paraphrase” of s2s_{2}’s meaning is a standalone proposition, and can be evaluated using the AIS definition for standalone propositions.

We will make this notion of “paraphrase” of the meaning of an utterance in context more formal, through the introduction of explicatures. The explicature of s2s_{2} in context of the previous utterances u1,s1,u2u_{1},s_{1},u_{2} is ee = George Harrison was 25 years old when his album “Wonderwall Music” was released. Once explicatures have been defined in this way, they can be evaluated for AIS in exactly the same way as standalone propositions.

We will use the following definition of interaction throughout the paper:

An interaction consists of: 1) a sequence u1…umu_{1}\ldots u_{m} of m≥0m\geq 0 user utterances; 2) a sequence s1…sns_{1}\ldots s_{n} of n≥0n\geq 0 system utterances; 3) a strict total order over the m+nm+n user and system utterances.For example, the order might be specified by functions U:{1…m}→{1…(m+n)}U:\{1\ldots m\}\rightarrow\{1\ldots(m+n)\} and S:{1…n}…(m+n)}S:\{1\ldots n\}\ldots(m+n)\} where U(i)U(i) (respectively S(i)S(i)) is the position of utterance uiu_{i} (respectively sis_{i}) in the total ordering. The notational details will not be important for this paper.

This setting is intended to be quite general, including a broad class of applications where systems generate utterances. In conversational QA systems we typically have alternating user and system utterances, where m=nm=n, and the total ordering is u1,s1,u2,s2,…un,snu_{1},s_{1},u_{2},s_{2},\ldots u_{n},s_{n}. In summarization tasks we have a simplified setting where m=0m=0, n=1n=1, and s1s_{1} is equal to the summary generated by the system. Table-to-text tasks are similar to summarization in that m=0m=0, n=1n=1, while s1s_{1} is the description of the table generated by the system.

Each sentence has an associated linguistic context:

We define the linguistic context for system sentence si,js_{i,j} to be ci,jc_{i,j}, where ci,jc_{i,j} is the ordered sequence of sentences (with speaker identities, user or system) that precedes si,js_{i,j} in the total ordering. We define the linguistic context for user sentence ui,ju_{i,j} to be ci,j′c^{\prime}_{i,j}, where c′c^{\prime} is defined in a similar way.An equally plausible definition would be to define ci,jc_{i,j} to also include the following sentences within utterance sis_{i}, that is, si,j−1,si,j+1…si,∣si∣s_{i,j-1},s_{i,j+1}\ldots s_{i,|s_{i}|} (and an analogous definition for ci,j′c^{\prime}_{i,j}). That is, the context would be extended to include sentences that follow si,js_{i,j} in the utterance sis_{i}. This would allow instances of cataphora, for example, to be handled in the definitions of explicatures and attribution.

Here the definition of “sentence” is intended to be quite broad. A sentence could be a declarative sentence, a question, or a fragment (such as the string “25 years old”). Under the above definitions, the context for a user or system sentence is simply the sequence of user and system sentences that precedes it. To illustrate these definitions consider the following example:

Here the system utterance s2s_{2} consists of two sentences, s2,1=s_{2,1}= He was 25 years old and s2,2=s_{2,2}= It was the first solo album by a member of the Beatles.

3 Explicatures

A key goal in this section is to define AIS for sentences si,js_{i,j} in linguistic contexts ci,jc_{i,j} which are non-empty (i.e., which contain previous sentences in the discourse). To do this it will be critical to formalize what is meant intuitively by “the meaning of si,js_{i,j} in context ci,jc_{i,j}”. To do this we introduce explicatures (this definition is closely related to definition 1 in Choi et al. (2021)):

Define the context cc to be (cl,t)(c_{l},t), where clc_{l} is the linguistic context and tt is the time. Define cˉ\bar{c} to be the context (ϵ,t)(\epsilon,t) where ϵ\epsilon is the linguistically empty context: that is, cˉ\bar{c} is a copy of cc but with clc_{l} replaced by ϵ\epsilon. The set of explicatures E(c,x)E(c,x) of a sentence xx in a context cc is a set that satisfies the following conditions: 1) each e∈E(c,x)e\in E(c,x) is a declarative sentence or question that is interpretable in context cˉ\bar{c}; 2) each e∈E(c,x)e\in E(c,x) has the same truth-conditional meaning in cˉ\bar{c} as the meaning of sentence xx in context cc.

Note that the sentence xx will most often in this paper be a system sentence si,js_{i,j} in linguistic context ci,jc_{i,j}, but can also be a user sentence ui,ju_{i,j} in linguistic context ci,j′c^{\prime}_{i,j}.

Thus, each e∈E(c,x)e\in E(c,x) is a paraphrase of xx that is interpretable in the linguistically empty context and that preserves the truth-conditional meaning of xx in context cc. Note that E(c,x)E(c,x) is a set because there may be multiple ways of paraphrasing xx, which are equivalent in meaning. Given an equivalence relation between sentences that identifies whether any two sentences are equal in meaning or not, we can think of a single member of E(c,x)E(c,x) as a representative of the entire set E(c,x)E(c,x). Following this, in a slight abuse of terminology we will henceforth often write “the explicature of xx in context cc is ee” as if there is a single unique explicature ee, with the understanding that ee represents the entire set E(c,x)E(c,x). We will also write E(c,x)=eE(c,x)=e as shorthand for E(c,x)E(c,x) being equal to the set of all sentences whose meaning is the same as that of ee.

In addition, we define interpretability as follows:

A sentence xx in context cc is uninterpretable if the truth-conditional meaning of xx in context cc is unclear. In this case we write E(c,x)=\mboxNULLE(c,x)=\mbox{\tt NULL}.

Figure 1 shows several examples illustrating these definitions. Some key points are as follows:

Remark 1: In example E1, the system response is a direct answer to a question, s2,1=s_{2,1}= 25 years old. s2,1s_{2,1} itself is not a declarative sentence, but given the context (in particular the question it is answering), its explicature is the standalone proposition George Harrison was 25 years old when “Wonderwall Music” was released. This type of example — where a direct answer to a question is an entity, noun-phrase, or some other fragment, but its explicature is a standalone proposition — is important and frequent. As another example consider the following:

Remark 2: In Example E3, the system segment is a sequence of two declarative sentences. Each sentence has an explicature that is a standalone proposition. This type of case is again frequent and important.

Remark 3: In Example E4 the system utterance is uninterpretable, because it is not clear what “the band” is referring to. Example E5 contains disfluencies that make it difficult to reliably interpret: “it” is not the expected pronominal reference; in this context “25” becomes too ambiguous to interpret as referring to the age of a human entity.

Remark 4: Examples E6 and E7 contain questions in the system and user utterance respectively. These examples illustrate that single questions (E7) or questions within multi-sentence utterances (E6) have well-defined explicatures.

With this background, we can now give the full definition of AIS:

A pair (s,c)(s,c), where ss is a sentence and c=(cl,t)c=(c_{l},t) is a pair consisting of a linguistic context and a time, is Attributable to Identified Sources (AIS) iff the following conditions hold:

The system provides a set of parts PP of some underlying corpus KK, along with ss.

ss in the context cc is interpretable (i.e., E(c,s)≠\mboxNULLE(c,s)\neq\mbox{\tt NULL}).

The explicature E(c,s)E(c,s) is a standalone proposition.

The pair (E(c,s),t)(E(c,s),t) is attributable to PP.

The pair (E(c,s),t)(E(c,s),t) is attributable to a set of parts PP of some underlying corpus KK iff: A generic hearer will, with a chosen level of confidence, affirm the following statement: “According to PP, E(c,s)E(c,s)”, where E(c,s)E(c,s) is interpreted relative to time tt.

The definition is very similar to the earlier definition of AIS for standalone propositions, but with checks for interpretability, and with attribution applied to explicatures of system sentences. Note that AIS can only hold for system sentences that have an explicature that is a standalone proposition (condition 3). For example, the explicature in Example E6 in Figure 1 is not a standalone proposition, as it is a question. We leave the treatment of cases such as these to future work (we might for example evaluate attribution for declarative sentences within the explicature, excluding questions; or we might evaluate presuppositions within the questions themselves).

3.2 Attribution of Entire Utterances

In the previous sections we have described AIS for the individual sentences si,1…si,∣si∣s_{i,1}\ldots s_{i,|s_{i}|} within a system utterance sis_{i}. This assumes that such a segmentation of the utterance into sentences is available, for example, it is provided by the system. An alternative is to evaluate entire utterances sis_{i} for AIS, in a “single-shot” annotation. AIS applied at the utterance level could potentially have the advantages of simplicity, and the avoidance of segmenting utterances into sentence boundaries. It has the potential disadvantage of being coarser grained, not allowing AIS judgments at the sentence level. The choice of sentence-level vs. utterance-level AIS will depend on the exact application of AIS.

It should be relatively straightforward to extend the full definition of AIS (Section 3.3.1) to apply to multi-sentence utterances. The definition of explicatures would need to be extended to multi-sentence utterances; the definition of standalone propositions would also have to be extended to apply to multiple sentences; the definition of “attributable” would also need to be extended.

4 Towards Operationalization of AIS

In the above definition of AIS, three definitions are of key importance: 1) the “according to” test for standalone propositions; 2) the definition of interpretability; 3) the definition of explicatures, which are related to the interpretation of utterances in non-empty linguistic contexts. Note that it is not necessary for annotators to explicitly wield all of these definitions, or come to understand any of them in entirely formal terms, in order to provide AIS judgments. In developing human annotation guidelines for annotators who do not necessarily have background in these concepts, we relay the “according to” test and interpretability in a way that leverages natural speaker intuitions. We convey explicature through the more intuitive idea of a sentence paraphrase with respect to linguistic context. Annotators are instructed to apply the “according to” test strictly, without making further assumptions beyond what is conveyed in the text.

The formal definition of AIS makes several idealizing assumptions that must be relaxed in practical settings. In lieu of the posited “generic hearer”, the judgments of actual annotators will naturally be influenced by the particulars of their interpretive capacities, stemming from differences, for example, in cultural background and domain expertise. Table 1 lists several instances where such differences could conceivably affect judgments. These effects are, to some extent, inherent to implementing AIS using human judgments.

Human Evaluation Study

We evaluate the feasibility of human AIS assessment for three NLG tasks: conversational question answering, summarization, and table-to-text generation. To quantify the significance of human judgements, we present evaluators with the output of different models for each of the tasks.

The set-up for these annotation tasks is to ask annotators to rate the AIS quality of ss, some model produced output given some attributed source PP. In the conversational QA and summarization settings, PP is a document or passage from a document, while in the table-to-text setting PP is a table and its description. For conversational QA, annotators are also provided with a context cc, which is the set of previous conversation turns. cc is used to help annotators understand the contextualized meaning of the model output, what we formally define as explicature in §3.3.

Because this is a challenging task with many possible edge cases (such as those discussed in Table 1), we ask five annotators to judge each example. In our results section, we compare to the consensus answer (if there is one) for simplicity. In future work, researchers who wish to use AIS for evaluating systems might find use in distinguishing between cases that are more clear cut (i.e., unanimous) versus those where there may be some inherent ambiguity.

We break the annotation task into two stages described in §4.1.1 and §4.1.2, which mirrors the formal steps in the AIS definition (§3.3.1) First, the annotators are asked if they are able to understand and identify the information being shared in the model output without seeing the source document (i.e., whether it is interpretable on its own). Then, if the output is deemed interpretable, the annotators are shown the “attributed source” P and asked whether all of the information that is shared in SS can be attributed to PP (i.e., whether it is AIS). As described in the results sections, the splitting of the task into these two steps helps annotators to first filter out outputs that are badly formed (e.g., ungrammatical to the point of impeded intelligibility) or too ambiguous (e.g., unclear pronouns) to appropriately evaluate the attribution. In the results, we report scores based on the annotator consensus (i.e., majority vote): the percent of total examples marked as interpretable (Int in Tables) and the percent of interpretable examples that were marked as AIS (AIS). In some datasets, certain examples were flagged as difficult to annotate due to legibility-related issues (see §4.1.3). For those cases, we separately report the percentage of examples that were flagged (Flag) and thus excluded from the interpretability and AIS scores.

In the initial stage of the annotation task, we show the annotators the model output ss and any preceding context cc without showing the source. We ask them to identify the interpretability by posing a yes/no question. For example, in the summarization task the annotators are asked:

Is all of the information relayed by the system summary interpretable to you?

Note that context cc is populated with preceding turns of the system–user interactionSome interactions may contain no previous turns. in the conversational QA task, whereas in summarization and table-to-text tasks it is always empty. In the instructions, the context cc is explicitly called out to be used in interpreting output ss in the conversational QA task.

Here, the goal is to tease out if the model-generated output ss contains any potential ambiguity that would prevent or misguide establishing attribution to its source PP. Anaphora resolution is the main source of this type of ambiguity, where deictic elements do not have clear antecedents within ss or its context — for example, pronominal usage with an unclear or broken coreference chain or definite noun phrases as first mentions. Additionally, syntactic ambiguity or disfluency may also result in diminished interpretability of ss (see Examples E4, E5 in Figure 1).

We acknowledge potential anthropomorphizing effects on how annotators interpret the system output Gopnik and Wellman (1992). Because cooperative meaning co-construction between interlocutors is the default communicative strategy of inter-human interaction Grice (1975), when faced with ambiguities and slight discrepancies in the system output, annotators may be “forgiving” of diminished interpretability, especially if the underlying source is present and can help recover missing context.

In our experiments we have found that not presenting the source at this stage is crucial for ensuring that evaluators are strict in their assessment of interpretability of the system output (see Figures 3, 5, and 7 for how it was implemented in the task interface).

1.2 AIS Rating

If an annotator selects “yes” for the interpretability question, we show them the source PP and ask them whether all of the information relayed in the output ss can be supported by PP. For example, in the conversational QA task the annotators are asked:

Is all of the information provided by the system response fully supported by the source document?

Note that the PP for the conversational QA task is the retrieved document that serves as the source of the system output ss. In the summarization task PP is the original news article from which the summary in ss was derived. In the table-to-text task PP is the original table, highlighted cells, and table metadata (table title, section title, and section text) from which the textual table description is generated.

In the instructions, we tell annotators to first think about all of the information that is contained the output including: what’s directly stated in the output sentence verbatim as well as any explicatures that can be made from the output with respect to the context, such as inferring pronoun references from the conversational history.

Annotators are instructed to only mark output as attributable if it is clear that all parts can be directly inferred from the source. The instructions specifically call out to utilize the paraphrase test:

In determining this question, ask yourself whether it is accurate to say “the provided news article says…” or “according to the news article…” with the system summary following this phrase.

If the output is misrepresenting information from the source because it is misleadingly worded, missing important context, or even changing only slight details, these cases are all counted as “not fully attributable”.

1.3 Flag Rating

A special rating is reserved for flagging items that would be disqualified from the the task altogether because they flout the range of possible relationships between the utterance, its context, and the source defined in 3.3.1.

In practical terms, these are tasks that are too malformed for annotators to perform judgements on. This category includes tasks with rendering issues in the interface (missing task elements, e.g., empty utterance), corrupted text resulting in non-communicative utterances (bad text encoding, HTML artifacts), underspecified source (the source document itself is ambiguous because it is too short and may contain unresolved reference chains), or a source that is difficult to understand because it requires expert-level knowledge.

Once a task is flagged, it is disqualified from the rating queue of the annotator who flagged it. Other annotators may choose not to flag this item; cumulative ratings and interannotator agreements are calculated for all non-flagged ratings of a task (see the flag sections in the annotator guidelines for conversational QA, summarization, and table-to-text) Section Flag.

1.4 Limitations

By asking yes/no questions, we can greatly reduce the complexity of this task for annotators. However, for some applications of AIS measures, it may be useful to have more fine-grained measures. Additionally, we ask annotators to evaluate the entire output (rather than sentences or specific spans) under the reasoning that if even one span within the model output is not AIS, then the whole output is not AIS (cf. Maynez et al. (2020), Durmus, He, and Diab (2020)).

We also acknowledge that there are other aspects of model output quality (e.g., relevance, non-redundancy, etc.) not evaluated here. We focus on the separate evaluation of AIS as part of a focused effort towards quantifying the attribution itself, disentangled from other desirable generation qualities.

2 Human Evaluation Procedure

The ratings were performed by a group of nine paid full-time annotators under the guidance and supervision of a project manager. The annotator team is based in Hyderabad, India; the annotators are native speakers of the Indian dialect of the English language. The annotators do not have a background in linguistics. They were trained for this specific task.

Three separate user interfaces were developed for performing the evaluation in this study: one for the conversational QA tasks evaluating the output of models trained on QReCC and WoW datasets, another for summarization tasks evaluating the output of models trained on the CNN/DM dataset and lastly one for table-to-text tasks evaluating the output of models trained on the ToTTo dataset. The interfaces share many fundamental design elements with task dependent modifications. For example, the conversation QA interface contains a devoted element for displaying the conversational history. All three interfaces explicitly hide the source document/table at the stage when interpretability of the system output is evaluated (see the Appendix for the interface layouts and annotator prompts (Figures 3, 4, 5, 6, 7, and 8).

The annotators were trained on the tasks in a series of stages. First, a pilot study of 50–100 items was conducted with the first iteration of the annotator instructions. As part of the pilot, all ratings were required to have written justifications elaborating the reasoning for the provided rating. The results of the pilot were analyzed by the authors to identify common errors patterns; collected justifications were helpful in understanding the reasoning annotators used to arrive at their ratings. The results of the review were communicated back to the annotators, and the instructions were modified to emphasize areas leading to common ratings errors.

Next, a portion of the ratings was inspected by the authors for persistent error patterns and the feedback communicated to annotators. Additionally, the annotators collected edge cases where they found it difficult to make judgements. These edge cases were adjudicated by the authors; recurring complex patterns were used to expand the annotator guidelines (see the Appendix for full final instructions for conversational QA, summarization, and table-to-text).

Finally, the annotator team performed internal audits on a subset of completed tasks.

Annotators were initially trained on the conversational QA tasks; other tasks and training were introduced subsequently.

Experiments

In the following section, we demonstrate the utility of the AIS templates by showing how it can be applied to three different tasks (conversational QA, summarization, and table-to-text generation) in which the model output is — by design — always meant to be attributable to some source document. We instantiated the AIS annotation template for four datasets in these domains (see Table 2) and performed human evaluation studies on generated outputs from multiple models. In order to show the applicability of AIS in detecting nuanced differences between different types of model outputs, we specifically chose models for each dataset that would represent a range of different types of outputs rather than just selecting a set of state-of-the-art models. We also annotated a selection of gold references from each dataset to better understand the AIS quality of existing datasets in these areas. We end with analysis of how effectively humans can annotate AIS as well as a discussion of various interpretability and AIS patterns that we found in the resulting annotations.

We use the QReCC dataset Anantha et al. (2021), a collection of multi-turn conversational QA interactions that extends conversations coming from NaturalQuestions Kwiatkowski et al. (2019), QUAC Choi et al. (2018), and CAST-19 Dalton et al. (2020). In this task, a model is given a conversational history and generates a contextualized response. We use a task set-up where the document passage containing the answer to the current query has already been retrieved (using the oracle retrieved document passage as the attributed source). We use different variations of T5 models including both base and small size variants. First, we use the pre-trained T5 models (PT) by themselves by prompting the model (formatted as: “Query:… Conversation History: … Document: … Answer:”). We also use a version of T5 that has been fine-tuned on QReCC (FT) which uses special tokens to separate the query, context, and document instead of natural-language prompts. Lastly, to sanity-check the AIS measures, we use a version of the model (no evidence) that only sees the query and conversation history but not the document at generation time. We expect that the AIS subscores should be much lower in the model that does not use the evidence from document to generate the answer.

We show results in Table 3. The model outputs’ interpretability increases substantially after fine-tuning (by about 50 points). The AIS subscore is highest in the fine-tuned model that uses evidence in its input. As expected, the AIS is drastically lower in the model that does not use the document as input at generation time (the no evidence model) which is both interpretable and AIS only 15% of the time. Differences between model sizes (small vs. base) are generally not significant except for the pretrained-only model, though the AIS scores of the smaller versions are typically slightly higher.

2 WoW Answer Generation

We used the seen portion of the test set from Wizard of Wikipedia Dinan et al. (2019). In this task, a model is given a conversational history and generates a contextualized response based on information from Wikipedia. As with QReCC, we again use a set-up where the Wikipedia sentence has already been retrieved (using the oracle retrieved sentence as the attributed source). To avoid chit–chat style utterances that may not be sharing new information, we sampled 200 examples per model where the previous utterance was a question (contains ‘?’). We used the models from Rashkin et al. (2021). That paper introduced a controlled T5 model trained on the Wizard of Wikipedia data which uses control tags and re-sampling to target generations that are more faithful to the document (by looking at heuristics such as entailment metrics, lexical precision, and first-person usage). Similar to that paper, we also compared with three models that are seq2seq-style conversation models: the original answer generation system from Dinan et al. (2019), the Dodecadialogue multitask system from Shuster et al. (2020) and a T5-base model (Raffel et al., 2020) finetuned on Wizard of Wikipedia data. Because the model from Rashkin et al. (2021) was specifically trained to be more faithful to evidence, we expect that it will score higher in the AIS category.

We show results in Table 4. Compared to the QReCC data (in which only a few examples were flagged), more examples were flagged with the Wizard of Wikipedia data, which we included as an extra column. The general trend of results is similar to what was found in the human evaluations of faithfulness and subjectivity in Rashkin et al. (2021). As expected, the model that has specific controllable inputs for increasing the model’s faithfulness to the input document achieves the highest the AIS scores overall. We also note that the AIS scores of the gold references is lower than the model outputs. We discuss this more in Section 5.5.5.

3 CNN/DM Summarization

We extend our evaluation framework for a second task, summarization to confirm that AIS can be more broadly applicable. AIS is crucial in summarization where a generated summary (SS) must be well-supported by the source article (PP). In contrast to some of the prior work in hallucination evaluation in summarization (Durmus, He, and Diab, 2020; Maynez et al., 2020), the annotators in our task evaluate the full summary for attribution (rather than at a sentence-level or a span-level), in order to account for cases where two individual text spans may be attributable to a source document but — when composed together — convey information that is different from the source document (e.g., misordered events, pronouns that no longer have the correct references when misordered, etc.). As a first step in applying AIS to summarization, we compare the performance of three different approaches (abstractive vs. extractive vs. hybrid) on 200 examples randomly sampled from the CNN/DM (Nallapati et al., 2016) test set. The source articles in this dataset come from articles in CNN and DailyMail news and the summaries are extracted from bulleted highlights that were included with the article by the journalists. We expect that high-quality AIS annotations will show a trend where extractive systems achieve higher AIS scores because they are copying directly from the source without adding anything. First, we used MatchSum (Zhong et al., 2020), a state-of-the-art extractive summarization model. Because this model is extractive, it is expected that it will be the least prone to hallucinations. We also used an abstractive summarization system, BigBird (Zaheer et al., 2020). Lastly, we used Pointer-generator Networks from See, Liu, and Manning (2017) — a hybrid approach that is uses an abstractive seq2seq model but with an explicit copy mechanism that can extract information from the source document.

We show results in Table 5. The more extractive approaches generally reach higher AIS subscores. This is a somewhat expected result — extractive systems are less likely to output hallucinations as they are quoting information verbatim from the documents. As with Wizard of Wikipedia, the AIS scores of the gold reference summaries is surprisingly lower than the model output, which we will discuss more in Section 5.5.5.

4 Table-to-Text ToTTo data

Lastly, we show the utility of extending AIS to a table-to-text task where PP is a table rather than a text document. SS is a sentence generated by a model to describe some highlighted portion of the table. We chose the ToTTo dataset Parikh et al. (2020), testing with T5 and ByT5 models that were previously used with this data in the GEM benchmark Gehrmann et al. (2021). We experiment with two different sizes of ByT5 and three different sizes of the T5 architecture. As before, we sampled the output of 200 examples from the test set. We also annotated 200 ground-truth references from examples in the dev. set (as the test set does not have gold-truth references publicly available).

We show results in Table 6. The model with the most “interpretable” responses was T5-base, with the ByT5 architectures being significantly less interpretable. On the other hand, the T5 architecture responses were more likely to be flagged (according to the annotators this was because the flagged responses contained artefacts like unintelligible character encoding errors). Generally, we don’t observe statistically significant differences in the AIS subscores though the larger architectures tended to have slightly lower AIS scores (similar to our observations of Table 3).

5 Annotation Quality

In this section we discuss the further implications of the human annotation results. We focus on two primary questions: (1) can humans reliably annotate AIS? and (2) what do our measured AIS ratings indicate about NLP data and models?

We show the interannotator agreement (IAA) for crowd annotators in the left half of Table 7. The metrics we used include Krippendorff’s alpha comparing individual ratings, pairwise agreement (PA) comparing individual ratings and an F1 score comparing individual ratings to the consensus (majority vote). Agreement results are generally moderate to high, displaying that — while this is a challenging task — the annotators are able to be fairly consistent with one another. The alpha scores are generally lowest on the summarization CNN/DM task, perhaps because the output text is much longer in summarization, increasing the complexity of the rating task. The F1 scores are similarly high, particularly on the AIS ratings.

5.2 Audits

Separately, the annotator team also performed internal audits on the annotation quality where a project lead from the annotator team examined a sample of individual annotator judgements at different points (snapshots) of the annotation process (Table 8). QRECC and WoW annotations were evaluated together as the broader conversational QA annotation task. The overall reported quality is in the high nineties for all three tasks with slight variations. The annotation quality for the conversational QA tasks remains high across all snapshots; we attribute this to the annotators extended experience with the task prior to the annotation of this datasetThe annotator pool was involved in annotating a series of related tasks for Conversational QA beyond the reported results in this paper.. The quality of the summarization annotations shows an increase over snapshots, as annotators internalize the guidelines and gain expertise in the task. The quality of the table-to-text annotations fluctuates and is generally the lowest of the three tasks; we attribute this to a much larger sample for which quality was measured. Overall, across the three tasks, the larger the quality evaluated sample, the lower the overall reported quality. Barring genuine task differences that would lead to variations in annotation quality, this suggests that the reported table-to-text quality of annotations is the most representative of all three tasks.

5.3 Crowd Annotator Performance

We also observe average task completion times decrease across all types of tasks as the annotators are exposed to more tasks and internalize the instructions (Table 9). Initial pilots for the conversational QA, summarization, and table-to-text tasks included required rating justifications for interpretability and AIS questions as part of annotator training. Once the production annotation started and the justifications were no longer required, completion times decreased significantly for all types of tasks, which we attribute primarily to the annotators no longer typing out detailed justifications of their ratings, but also overall internalization of the guidelines. The effects of annotators internalizing the guidelines are also evident when comparing completion times at production annotation start and finish: average task completion times decrease for all three types of tasks as annotators are exposed to more tasks and gain more experience.

At the same time, the absolute task completion times are consistently and substantially different across the three tasks suggesting their uneven complexity, with conversational QA taking the shortest amount of time to complete, summarization requiring the longest, and table-to-text falling in-between. This pattern follows the trend in the distribution of inter-annotator agreement across the three tasks: tasks with shorter completion times generally have higher interannotator agreement. We postulate that this is primarily due to the difference in the amount of context that is necessary to perform ratings. Although conversational QA tasks may contain several turns of preceding interactions between the system and the user as well as the source document, the amount of information in the source articles in the summarization task is substantially larger. Likewise, source tables in the table-to-text task can be extensive and have the added information complexity of cell highlighting and table metadata. Finally, register and discourse structure effects may be at play here as well. Conversational QA tasks build upon colloquial interactions between the user and the system, setting up the context of the interaction in shorter utterances and helping annotators anticipate the contents of the source document. Likewise, Wikipedia, news articles, and tables package information differently as they serve somewhat different communicative goals, and it is possible that one of these source types is more amenable to inspection required for performing AIS ratings.

5.4 Expert Ratings

Where AIS is used as a metric for ranking generative models, the internal consistency of crowd annotations is paramount. But, to help illuminate the inherent challenges in calibrating this annotation task, we also compare the crowd ratings with those of expert on a small set of examples. Due to the challenges of scaling expert evaluations, we limited expert ratings to two tasks (CNN/DM and QReCC) with 50 examples each. The experts (two co-authors) first annotated the examples separately from each other using the same interface as the crowd annotators and then discussed their answers to reach a consensus. Expertise here might be derived from general educational background (a different approach to close reading), the ability to discuss annotations (and to do so carefully at self-guided pace), specialized knowledge, and first-hand familiarity with the evaluation framework. Expertise does not imply that the experts have more experience performing the task than the crowd annotators.

In order to account for natural ambiguity in assigning a rating category, experts marked some cases as “either option acceptable”. We compare the individual crowd annotator ratings to the expert consensus in the right half of Table 7. Crowd annotators tend to agree with each other more than they agree with experts, which is expected due to differences in background, incentives, and procedure, although there is still reasonably consistent agreement in most cases. On closer inspection, we find that most disagreements are cases where there is underlying ambiguity caused by vagueness in the evidence or model output. In these cases experts erred more on the side of being critical of the model and crowd annotators erred more towards being lenient. In the case of conversational QA, most of the AIS disagreements involved cases where the document and the response do not refer to an entity using the same naming conventions (e.g., using both first and last name; see Table 10) leaving some ambiguity that the document is referring to the same entity as the response. The greatest source of disagreements overall is the interpretability question in the summarization task (see examples in Table 11). The summaries in the CNN/DM dataset were originally crawled from high-level article highlights, and experts observed that — due to the linguistic style of these highlights — there were many cases where the language may be vague or ambiguous, making this dimension more challenging. Because we use interpretability as a pre-filtering stage for the AIS question, we make allowances for the annotators being more inclusive. Despite the differences on the interpretability dimension, they generally agreed with experts on most AIS questions, our primary evaluation dimension.

5.5 Limitations of Gold References

The last rows in Tables 3, 4, 5 and 6 show annotation results on reference answers sampled from these datasets. The results demonstrate that there is actually a limit on the AIS quality of the data itself in multiple tasks. We include examples of non-AIS references in Table 12 to illustrate what some of these examples look like. We hypothesize that this is because the originators of the data were not specifically instructed to be as faithful to the underlying documents as possible. In the case of Wizard of Wikipedia (Dinan et al., 2019), the gold response is only AIS 16% of the time. But, this dataset was constructed for a different objective — to contain both informative and engaging responses. The MTurk workers who created the data were provided documents to enhance their conversations but could do so at their own discretion, often including their own thoughts and opinions in the conversation as well. This is also reflected in the CNN/DM AIS scores — summaries in CNN/DM are only attributable to the documents in 54% of the interpretable examples. Looking more closely, we speculate that this may be due to the post-hoc data creation process used to extract summaries from article highlights written by journalists. We observed that the reference summaries in CNN/DM may sometimes refer to external pieces of information that may have accompanied the article (a picture, a headline, etc.) or sometimes make assumptions about what the intended audience of the article might already know that can affect either the interpretability or AIS scores (see Example 1 in Table 12 and Example 3 in Table 13). These results indicate that there is still a need for high-quality AIS data for training new NLG models.

5.6 Examples

In the Appendix, we separately list textual examples rated as uninterpretable (Table 13), interpretable but not AIS (Table 14), or both interpretable and AIS (Table 15). For the table-to-text task, we present examples in a more visual figure, Figure 2, for better legibility. Common factors in marking text as “uninterpretable” include repetitive, degenerate language and ambiguous pronouns and ellipses. Additionally, some outputs are marked as uninterpretable because they are hard to understand “on their own”. Whether or not a piece of text can be understood may also rely on things like commonsense and background knowledge that could vary depending on annotators’ backgrounds (see Example 3 from Table 13). Ambiguous references can also affect both interpretability and the AIS scores. In Example 2 of Table 14, the retrieved document did not provide enough information to completely verify the response since it never refers to Ann Veneman by her full name. This is a seemingly minor detail, but annotators were often sensitive to this type of example since they could not verify whether the document was actually referring to the same entity as the model output. Another type of non-AIS output that frequently appeared in the QReCC data were cases where a model outputted a seemingly informative statement that — instead of being grounded to the document — was actually grounded to a previous conversation turn, sometimes repeating itself verbatim. Lastly, examples verify that AIS evaluations can be disentangled from other quality aspects, such as conversational relevance. This was challenging to instruct to annotators as it is instinctual to judge quality more holistically, and they were explicitly given instructions with multiple examples illustrating what types of quality aspects to ignore. In the resulting annotations, they would mark incoherent summaries or irrelevant conversational replies as AIS if they conveyed well supported information, appropriately disregarding other aspects of quality.

Discussion

Generative models have been advancing toward human-like competence in some aspects. Their real-world application in consumer-focused information products are becoming more attractive, for example, for summarizing original descriptions of events, or for deriving answers to pertinent questions about the world. Traditionally, this type of information transformation has been performed by specialized human experts (e.g., journalists, researchers), who are required to meet a variety of standards of accuracy and accountability, maintaining one or more sources for a proposition and performing fact-checking. The task could also be likened to the practice of law, where norms are examined for their subsumptive relationship to a set of circumstances, and where both close reading and a set of conventionalized tests aid this determination.

We formalize a specific sub-task of fact-checking, namely, verification against a known source, as a necessary but not sufficient step in ensuring the quality of generated text. We show that with the right training, careful instructions, and optimized user interfaces, we can delegate the judgment of attribution to underlying source(s) to crowd workers, but we also find limitations. Following the data collection we described, we found it necessary to set some standards in our instructions to raters. This includes setting expectations for named entities, for example, whether first and last names are needed to identify an individual and to link them between evidence and statement, or if a place name without qualification may be acceptable as long as there are no other well-known places of the same name. Similarly, as statements and evidence become more complex, raters inevitably draw inferences using individual world knowledge. This is unavoidable and is inherently noisy Pavlick and Kwiatkowski (2019); ground truth is ambiguous, just like journalists, researchers, or judges often legitimately disagree. Possible model outputs fall on a spectrum ranging from synthesized information to the mostly unassailable extractive generations Ladhak et al. (2022). AIS does not set policy about where model output should fall: its users still need to decide where to draw the line.

AIS is limited to propositions that can be judged with the “according to” framework. AIS is not applicable to questions (without presuppositions) or imperatives (commands and requests). There are also scenarios where strict attribution contradicts other desirable output characteristics (e.g., chit–chat systems). We did not examine AIS on such data. How to evaluate hybrid systems that mix entertaining and informative communicative goals — capturing the attribution of the informative portion but ignoring the rest — is unclear, as is the question of whether systems with blurry boundaries between what is and is not subject to attribution should exist at all.

We have purposefully limited the availability of context in our definition. Practical human–computer interactions may actually take place in context beyond the shared time tt that is used in the definition (Section 3), perhaps because the communication channel is richer than a text-based line of transmission, and because it may be further extended by multi-session interaction history. It is important that annotators remain aware of the notion of explicature, resolving explicit references and implicit topics available to the communicators. It is possible that the use of models that perform this task Choi et al. (2021) can improve the performance of raters. We are also aware that this task requires close reading, which is challenging to implement on crowdsourcing platforms where speed, efficiency, and cost are incentivized instead. Again, models may be useful in extracting explicit, elementary propositions from complex statements, making this task easier for raters. We will examine such approaches in future work.

Conclusion

In this paper, we define a new evaluation framework called Attributable to Identified Sources which allows us to inspect whether information in generated text can be supported by source documents. We provide formal definitions of AIS and descriptions of how it can be applied to three different NLG tasks (conversational QA, summarization, and table-to-text generation). We validate this evaluation framework quantitatively on human evaluation studies, in which annotators rated the AIS of model output as part of a two-stage annotation pipeline. The results of the human evaluation studies demonstrate that high-quality AIS ratings can be obtained empirically. The results shed light on some of the ongoing challenges in training NLG models; having solid AIS is the basis for addressing them.

Evaluation Instructions for Conversational Question Answering

The following is a verbatim representation of the instructions that were presented to paid crowd annotators for performing the task alongside the interface. The prompts in the rating interface include wording from the instructions; the rating interface also contains hyperlinks to example sections in the instructions for each question and rating.

Overview

In this task you will evaluate the quality of a system-generated response to a user query. The system is trying to help the user learn about a particular topic by answering their questions. We want to rate the system response quality based on how well it represents the original source.

We will be using two categories to evaluate the quality of the summary: Interpretability and Attribution. You will evaluate these categories in succession. Some ratings will result in other categories being skipped. The task interface will guide you through the flow; you can also see the overall task flow in the diagram below.

Note: The system-generated responses may appear very fluent and well-formed, but contain slight inaccuracies that are not easy to discern at first glance. Pay close attention to the text. Read it carefully as you would when proofreading.

The sections below describe each of the dimensions in detail. You can also flag ineligible tasks; the flagging criteria are described in this section.

Interpretability

In this step you will evaluate whether the system response is interpretable by you.

You will be shown an excerpt of a conversation between a user and an assistant-like computer system. The last turn in the conversation will be a user query from the user followed by a system response attempting to respond to their query. Given the context of the user query, carefully read the system response and answer the following question:

Is all of the information relayed by the system response interpretable to you?

This is asking whether you can understand the response. If there is some part of the response that is unclear or hard to interpret, select “No”. If prompted by the interface, enter a succinctly detailed justification of your rating.

An uninterpretable response has diminished intelligibility due to:

Vague or ambiguous meaning, e.g., unclear pronouns usage.

Malformed phrases and sentences that are difficult to understand.

If the system response is interpretable by you, you will proceed to the next category.

Examples of interpretability ratings and justifications are in this section.

Attribution

In this step, you will evaluate how well a system-generated response is attributable to the source document. Note that the source document is a new element that will appear in the task only when you reach this question.

Note: We refer to “attributable to source document” interchangeably as “attribution” and “supported by the source document”. By which we mean, all of the information in the system response can be verified from the source document.

You will be shown an excerpt of a conversation between a user and an assistant-like computer system. The last turn in the conversation will be a user query from the user followed by a system response attempting to reply to their query. You will also be shown a document that was cited by the system as its source in attempting to answer the question (source document). You will use all three (user query, system response, source document) to answer the following question:

Is all of the information provided by the system response fully supported by the source document?

This is asking whether all of the information in the system response can be attributed to the information present in the source document. If prompted by the interface, enter a succinctly detailed justification of your rating.

Attribution of system-generated response in relation to the source document can be established by considering the following:

What is the information provided by the system response?

Is this information an accurate representation of information in the source document?

A. Definition of “the Information Provided by the System Response”

Two points are key in determining the information provided by the system response:

The context of the system response — that is, the query and previous conversation turns — is often critical in determining the “information provided by the system response”.

The source document should be completely ignored when determining “the information provided by the system response.” (i.e., it should not be used as additional context).

In the above example, the meaning of the system response is clear even without seeing the query. But consider another example:

In this case the pronoun “he” depends on the context (i.e., the query): but it is clear that the intended meaning of the system response can be paraphrased as something along the lines of “Doug Williams in days of our lives is played by Bill Hayes”. In this case this paraphrase is the “information provided by the system response”.

Pronouns such as he/she/it/they etc. are one case where context is needed to figure out the intended meaning of the system response. Other examples are the following (given with paraphrases of the information that is provided by the system response):

In system response 1, the phrase “the show” needs context for interpretation, but it is clear from the context of the query that it refers to “days of our lives”. In system response 2, the system gives a direct answer to the query, simply “NBC”, but it is clear given the query that the information provided by the system is “days of our lives is on NBC”. In query 3, the phrase “56 seasons” needs context for interpretation, but given the query it is clear that the response is referring to “56 seasons of days of our lives”.

In general, use your best judgment to determine the information provided by the system response. If you are unsure what the intended meaning is of the system response, make sure that you marked the example as “No, the response is unclear.” as part of the Interpretability stage. As one example, take the following:

In this case it is not clear what “it” is referring to, and the meaning should be marked as being unclear. Again, use your best judgment in determining whether or not the meaning of the system response is clear.

B. Definition of “An Accurate representation of Information in the Source Document”

Again, you should use your best judgment in determining whether all of the information provided by the system response is “an accurate representation of information in the source document”. We give the following guidance:

In determining this question, ask yourself whether it is accurate to say “the document says…” or “according to the document…” with the system response following this phrase. For example, is it accurate to say “according to the document below, In the American daytime drama Days of Our Lives, Doug Williams and Julie Williams are portrayed by Bill Hayes and Susan Seaforth Hayes” in the example given above?

Be sure to check all of the information in the response. If only some of the information is supported in the document, but other parts of the information are missing from the document or not an accurate representation, then please mark “No, not fully attributable.”

The concept of “accurate representation” should be close to a journalist’s conception of this phrase. For example take this excerpt from this page on Accuracy in the NPR Ethics Handbook: “When quoting or paraphrasing anyone…consider whether the source would agree with the interpretation…” In other words, if you had written the source document, consider whether you would view the system response as an accurate representation of information in that source document.

Some Final Important Notes

When making your judgments in this template, do not take into account whether the underlying source document is correct or trustworthy. This is clearly important, but will be evaluated in a separate task. The “attribution” category is used only to judge whether the information provided by the system response is an accurate representation of the underlying source document.

Additionally, when rating attribution, do not take into account the degree of relevance of the system response to the user query. Partially and fully relevant responses can be equally assessed for attribution, regardless of how much information contained in them is relevant to the query. In both cases you should judge whether or not the system responses are an accurate representation of information in the source document even if it doesn’t perfectly address the question.

Examples of attribution ratings and justifications are in this section.

Scoring and Examples

The response is unclear and/or difficult to understand.

Attribution

All of the information in the system response is supported by the document.

The response contains any amount of information that is not supported by the document (including responses that are only partially or not at all supported).

Flag

There is a flag button in the bottom left corner of the task interface. Once flagged, you can proceed onto the next task. Use it report tasks that are ineligible for reasons such as:

Some tasks may have missing user queries, responses or source text. They should be flagged.

Some text may be severely malformed with unintelligible artifacts (e.g. html code, unformatted tables, etc.). If any component of the task contains malformed text, the task should be flagged. The example below shows various types of malformed text.

Some source documents may not have sufficient information to determine whether it does/doesn’t support the response. They may be too short and/or lack critical information that would be necessary to rate the response.

Some documents may include scientific formulas, obscure terminology, etc. If you can still understand enough of the document to rate the attribution, please do so. But, on the other hand, if properly evaluating the response requires expertise in a particular area, please flag it.

Evaluation Instructions for Summarization

The following is a verbatim representation of the instructions that were presented to paid crowd annotators for performing the task alongside the interface. The prompts in the rating interface include wording from the instructions; the rating interface also contains hyperlinks to example sections in the instructions for each question and rating. The summarization instructions were developed after the conversational QA instructions had been established and the annotators had been trained on the conversation QA task.

Overview

In this task you will evaluate the quality of a system-generated summary. The system’s goal is to summarize the source news article, while remaining truthful to it. We want to rate the quality of the summary based on how well it represents the original source.

We will be using two categories to evaluate the quality of the summary: Interpretability and Attribution. You will evaluate these categories in succession. Some ratings will result in other categories being skipped. The task interface will guide you through the flow; you can also see the overall task flow in the diagram below.

Note: The system-generated summaries may appear very fluent and well-formed, but contain slight inaccuracies that are not easy to discern at first glance. Pay close attention to the text. Read it as carefully as you would when proofreading.

The sections below describe each of the dimensions in detail. You can also flag ineligible tasks; the flagging criteria are described in this section.

Interpretability

In this step you will evaluate whether the system summary is interpretable by you.

You will be shown a system-generated summary of a news article. Note that the news article from which the summary is derived is hidden at this step, because we need to evaluate whether the summary is interpretable on its own. Carefully read the summary and answer the following question:

Is all of the information relayed by the system summary interpretable to you?

This is asking whether you can understand the summary on its own. If there is any part of the summary that is unclear or hard to interpret, select “No”. If prompted by the interface, enter a succinctly detailed justification of your rating.

An uninterpretable summary has diminished intelligibility due to:

Vague or ambiguous meaning, e.g., unclear noun references or pronouns usage.

Malformed phrases and sentences that are difficult to understand.

If the summary is interpretable by you, you will proceed to the next category.

In the section below, we show in more detail the kind of reasoning that should be used for establishing interpretability of summaries. More examples of interpretability ratings and justifications are in this section of the appendix.

Interpreting the information provided in the system summary

In the above example, the meaning of the summary is clear even without seeing the original news article. It is clear what the summary is reporting on and it stands on its own, that is, this summary is interpretable. It should be marked as “Yes, I understand it.” But consider another example:

In this case the meaning of the phrase “the magnitude” obviously depends on some context, but that context is missing in the summary. Without additional information that clarifies that the magnitude refers to an earthquake that occurred in a specific location (Kathmandu), the summary is difficult to interpret and it does not stand on its own. It should be marked as “No, the summary is unclear.”

Noun phrases that require clarifications of this kind are one case where interpretability can be diminished. Other examples include nouns and pronouns without a (clear) reference and malformed phrases and sentences:

In summary 1, the phrase “the project” needs context for interpretation (“what project is being reported on?”). Likewise, “the museum” and “the city” are unclear (“what museum is this?”, “what city is this taking place in?”). In summary 2, the system provides a clear reference for the pronoun “he” (“hernandez”), but a reference for the pronoun “it” is missing (“what is scheduled to begin in may?”). Also it is not clear what “not legally required to get a conviction” is referring to. In summary 3, the first sentence is malformed because it is missing the announcement (“what did john stamos announce?”). The second sentence is difficult to understand (“who does ‘both’ refer to?”, “who will return to the new series?”).

In general, use your best judgment to determine the information provided by the summary. If you are unsure what the intended meaning of the summary is, err on the side of marking it with “No, the summary is unclear.”

Attribution

In this step, you will evaluate how well a system-generated summary is attributable to the source news article. Note that the source news article is a new element that will appear in the task only when you reach this question.

Note: We refer to “attributable to the source news article” interchangeably as “attribution” and “supported by the source news article”. By which we mean, all of the information in the system-generated summary can be verified from the source news article.

You will be shown a system-generated summary of a news article. You will also be shown the news article that was used by the system to generate this summary (source news article). You will use both of these to answer the following question:

Is all of the information provided by the system summary fully supported by the source document?

This is asking whether all of the information in the system summary can be attributed to the information present in the source news article. If prompted by the interface, enter a succinctly detailed justification of your rating.

a fully supported (or attributable) system-generated summary contains an accurate representation of information in the source news article. No information in the summary is unattested when compared against the source news article.

In the section below, we show in more detail the kind of reasoning that should be used for establishing attribution of summaries. More examples of attribution ratings and justifications are in this section of the appendix.

Assessing the accuracy of the information in the summary against the original news article

Again, you should use your best judgment in determining whether all of the information provided by the system summary is “an accurate representation of information in the source news article”. We give the following guidance:

In determining this question, ask yourself whether it is accurate to say “the provided news article says…” or “according to the news article…” with the system summary following this phrase.

Be sure to check all of the information in the summary. If only some of the information is supported in the news article, but other parts of the information are missing from the news article or not an accurate representation, then please mark “No, not fully attributable.”

The concept of “accurate representation” should be close to a journalist’s conception of this phrase. For example take this excerpt from this page on Accuracy in the NPR Ethics Handbook: “When quoting or paraphrasing anyone…consider whether the source would agree with the interpretation…” In other words, if you had written the source document, consider whether you would view the summary as an accurate representation of information in that source document.

Some Final Important Notes

When making your judgments in this template, do not take into account whether the underlying source news article is correct or trustworthy. This is clearly important, but will be evaluated in a separate task. The “attribution” category is used only to judge whether the information provided by the system summary is an accurate representation of the underlying source news article.

Examples of attribution ratings and justifications are in this section.

Scoring and Examples

The summary is unclear and/or difficult to understand.

Attribution

All of the information in the system summary is supported by the document.

The summary contains any amount of information that is not supported by the source new article (including summaries that are only partially or not at all supported).

Flag

There is a flag button in the bottom left corner of the task interface. Once flagged, you can proceed onto the next task. Use it to report tasks that are ineligible for reasons such as:

Some tasks may have missing summaries or news articles. They should be flagged.

Some text may be severely malformed with unintelligible artifacts (e.g. html code, unformatted tables, etc.). If any component of the task contains malformed text that makes it difficult for you to accomplish the task, the task should be flagged.

Some documents may include scientific formulas, obscure terminology, etc. If you can still understand enough of the document to rate the attributability, please do so. But, on the other hand, if properly evaluating the summary requires expertise in a particular area, please flag it. \appendixsectionEvaluation Instructions for Table-to-Text

The following is a verbatim representation of the instructions that were presented to paid crowd annotators for performing the task alongside the interface. The prompts in the rating interface include wording from the instructions; the rating interface also contains hyperlinks to example sections in the instructions for each question and rating. The table-to-text instructions were developed after the conversational QA and summarization instructions had been established and the annotators had been trained on the conversation QA and summarization tasks.

Overview

In this task you will evaluate the quality of a system-generated caption for highlighted parts of a table. The system is trying to convert the information in the table into a natural language description (what we are referring to as the “system-generated caption”). We want to rate the quality of the system-generated caption based on how well it represents information from the source table.

We will be using two categories to evaluate the quality of the caption: Interpretability and Attribution. You will evaluate these categories in succession. Some ratings will result in other categories being skipped. The task interface will guide you through the flow; you can also see the overall task flow in the diagram below.

Note: The system-generated captions may appear very fluent and well-formed, but contain slight inaccuracies that are not easy to discern at first glance. Pay close attention to the text. Read it carefully as you would when proofreading.

The sections below describe each of the dimensions in detail. You can also flag ineligible tasks; the flagging criteria are described in this section.

Interpretability

In this step you will evaluate whether the system caption is interpretable by you.

You will be shown a system-generated caption of a table. Note that the table from which the caption is derived is hidden at this step, because we need to evaluate whether the caption is interpretable on its own. Carefully read the caption and answer the following question:

Is all of the information relayed by the system caption interpretable to you?

This is asking whether you can understand the caption. If there is some part of the caption that is unclear or hard to interpret, select “No”. If prompted by the interface, enter a succinctly detailed justification of your rating.

An uninterpretable caption has diminished intelligibility due to:

Vague or ambiguous meaning, e.g., unclear noun references or insufficient context.

Malformed phrases and sentences that are difficult to understand.

If the system caption is interpretable by you, you will proceed to the next category.

In the section below, we show in more detail the kind of reasoning that should be used for establishing interpretability of captions. More examples of interpretability ratings and justifications are in this section of the appendix.

Interpreting the information provided in the system caption

In the above example, the meaning of the caption is clear even without seeing the source table. It is clear what the caption is reporting on and it stands on its own; that is, this caption is interpretable. It should be marked as “Yes, I understand it.” But consider another example:

In this case the meaning of the caption obviously depends on some context (“what did they finish?”), but that context is missing in the caption. Without additional information that clarifies that this is a result of a sports competition, the caption is difficult to interpret and it does not stand on its own. It should be marked as “No, the caption is unclear.”

Captions that require clarifications of this kind are one case where interpretability can be diminished. Other examples include nouns without a (clear) reference and malformed phrases and sentences:

In caption 1, the numbers are missing the context of what they are being reported on (“what location or event do these numbers represent?”). In caption 2, the phrase “a member” lacks a specifying reference (“what was Thomas Heflin a member of?”). In caption 3, the sentence is difficult to understand as it appears to be missing a noun (“was George A. Gilet a rugby player?”).

In general, use your best judgment to determine the information provided by the caption. If you are unsure what the intended meaning of the caption is, err on the side of marking it with “No, the caption is unclear.”

Attribution

In this step, you will evaluate how well a system-generated caption is attributable to the source table. Note that the source table is a new element that will appear in the task only when you reach this question.

Note: We refer to “attributable to source table” interchangeably as “attribution” and “supported by the source table”. By which we mean, all of the information in the system caption can be verified from the source table.

You will be shown a system-generated caption. You will also be shown a table and its associated descriptions: title, section title and table section text. These elements provide additional context for understanding the information in the table. Finally, some cells in the table will be highlighted as helpful hints for which parts of the table are the focus of the caption. The table, descriptions and highlighted cells (source table) were used by the system to create the caption. You will use all of these elements to answer the following question:

Is all of the information provided by the system caption fully supported by the source table?

This is asking whether all of the information in the system caption can be attributed to the information present in the source table. If prompted by the interface, enter a succinctly detailed justification of your rating.

a fully supported (or attributable) system-generated caption contains an accurate representation of information in the source table. No information in the caption is unattested when compared against the source table and its associated descriptions (title, section title and table section text).

In the section below, we show in more detail the kind of reasoning that should be used for establishing attribution of captions. More examples of attribution ratings and justifications are in this section of the appendix.

Assessing the accuracy of the information in the caption against the source table

Again, you should use your best judgment in determining whether all of the information provided by the system caption is “an accurate representation of information in the source table”. We give the following guidance:

In determining this question, ask yourself whether it is accurate to say “the provided table says…” or “according to the table…” with the system caption following this phrase.

Be sure to check all of the information in the caption. If only some of the information is supported in the table, but other parts of the information are missing from the table or not an accurate representation, then please mark “No, not fully attributable.”

The concept of “accurate representation” should be close to a journalist’s conception of this phrase. For example take this excerpt from this page on Accuracy in the NPR Ethics Handbook: “When quoting or paraphrasing anyone…consider whether the source would agree with the interpretation…” In other words, if you had written the source document, consider whether you would view the caption as an accurate representation of information in that source document.

Some Final Important Notes

When making your judgments in this template, do not take into account whether the underlying source table is correct or trustworthy. This is clearly important, but will be evaluated in a separate task. The “attribution” category is used only to judge whether the information provided by the system caption is an accurate representation of the underlying source table.

Some of the cells in the table are highlighted. The highlighted cells are intended to be the focus of the caption, and can be used as a helpful hint of where to look in the table for information in the caption — though some captions may also refer to information from elsewhere in the table. If the caption does not capture the information in the highlighted cells, but otherwise accurately represents the information elsewhere in the table and its description, please still mark it “Yes, fully attributable.”

Examples of attribution ratings and justifications are in this section.

Scoring and Examples

The caption is unclear and/or difficult to understand.

Attribution

All the information in the caption is supported by the table and its description.

The caption cannot be fully attributed to the source table and its description (including captions that are only PARTIALLY or NOT AT ALL supported).

Flag

There is a flag button in the bottom left corner of the task interface. Once flagged, you can proceed onto the next task. Use it to report tasks that are ineligible for reasons such as:

Some tasks may have missing summaries or news articles. They should be flagged.

Note that table title, section title, or table section text could be empty or designated with “None”. These are acceptable and should not be flagged. See an example of an acceptable table description below:

Some text may be severely malformed with unintelligible artifacts (e.g. html code, unformatted tables, etc.). If any component of the task contains malformed text, the task should be flagged.

Some tables may include scientific formulas, obscure terminology, etc. If you can still understand enough of the table to rate its attributability, please do so. But if properly evaluating the response requires expertise in a particular area, please flag it.

Annotator Interface for Conversational QA Tasks

Annotator Interface for Summarization Tasks

Annotator Interface for Table-to-Text Tasks

References