Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models

Bernd Bohnet, Vinh Q. Tran, Pat Verga, Roee Aharoni, Daniel Andor, Livio Baldini Soares, Massimiliano Ciaramita, Jacob Eisenstein, Kuzman Ganchev, Jonathan Herzig, Kai Hui, Tom Kwiatkowski, Ji Ma, Jianmo Ni, Lierni Sestorain Saralegui, Tal Schuster, William W. Cohen, Michael Collins, Dipanjan Das, Donald Metzler, Slav Petrov, Kellie Webster

Introduction

Large language models (LLMs) have shown impressive results across a variety of natural language understanding and generation tasks Devlin et al. (2019); Raffel et al. (2020); Brown et al. (2020); Rae et al. (2021); Zhang et al. (2022); Chowdhery et al. (2022); Chung et al. (2022) while requiring little or no direct supervision,By “direct supervision” we refer to labeled examples for the specific task in mind, for example datasets such as the Natural Questions corpus Kwiatkowski et al. (2019) for question answering. We use the term “direct supervision” to distinguish this form of supervision from the term “self supervision” sometimes used in the context of LLMs. instead using few-shot Brown et al. (2020) or in-context learning Xie et al. (2021). There is increasing evidence that LLMs may have potential in information-seeking scenarios, producing compelling output in scenarios ranging from “simple” question answering (e.g., Kwiatkowski et al. (2019); Rajpurkar et al. (2016); Joshi et al. (2017)), to long-form question answering Amplayo et al. (2022); Stelmakh et al. (2022), and information-seeking dialog Thoppilan et al. (2022); Glaese et al. (2022); Shuster et al. (2022); Nakano et al. (2021). This lack of direct supervision is particularly appealing given the difficulties of constructing labeled datasets for even simple question answering,Here we are referring to the traditional approach to data collection for supervised learning, where human raters provide labeled examples. An alternative approach is to use an LLM to generate labeled examples that are then rated by humans. For many tasks, this latter approach is considerably simpler. let alone more complex (but important) tasks such as multi-faceted question answering or interactive information-seeking dialog.

In many information-seeking scenarios, the ability of an LLM to attribute the text that it generates is likely to be crucial for both system developers and users (see Metzler et al. (2021); Rashkin et al. (2021); Menick et al. (2022); Thoppilan et al. (2022), and section 3.1, for a discussion). Ideally, an “attributed LLM” would seamlessly provide evidence snippets that support the text that it generates where appropriate (specifically, whenever it makes statements about the world, e.g., see Rashkin et al. (2021)). While there has been important work in the direction of adding attribution to LLMs (see Section 2), we argue that we as a field currently have very limited understanding of the challenge and how to make progress. Critical questions are:

How well do current state-of-the-art methods perform on attribution? Even for the simplest possible information-seeking scenario, simple QA, this is not well understood.

To explore these questions, we propose Attributed Question Answering (QA). In our formulation, the input to the model/system is a question, and the output is an (answer, attribution) pair where answer is an answer string, and attribution is a pointer into a fixed corpus, e.g. of paragraphs. The returned attribution should give supporting evidence for the answer; for example, it should satisfy the conditions in Rashkin et al. (2021) (see Section 3.1). Figure 1 gives an example.

Our motivation for studying attribution in QA is two-fold. First, it is perhaps the simplest information-seeking application, and as such it is more straightforward to evaluate. However, in spite of its simplicity, models and experiments for attributed QA are likely to be highly informative to the general goal of building attributed LLMs (see Section 3.1 for more discussion). Second, Attributed QA is an interesting task in its own right. It has advantages over existing approaches to evaluation of question answering systems (see Section 3.1 and Section 5). Attribution provided by a QA system is likely to be of benefit to both system developers and users. With this motivation, we make the following contributions.

First, we define a reproducible evaluation framework for Attributed QA, using human annotations as a gold standard. To facilitate progress, we additionally study AutoAIS Gao et al. (2022), an automatic metric that formulates evaluation as a Natural Language Inference task Dagan et al. (2005); Bowman et al. (2015). We find strong correlation between the two, making AutoAIS a suitable evaluation strategy in development settings.

Further, we perform a systematic analysis of a broad set of systems based on state-of-the-art components, exploring different architectures and levels of supervision. While retrieve-then-read architectures are attractive for their strong performance, they typically require a large amount of data to train and can be resource intensive. We are excited by the possibility of post-hoc attribution of LLM-generated answers (though this remains challenging), and end-to-end modeling that makes limited use of QA examples. We release scored system outputs to foster further exploration, at https://github.com/google-research-datasets/Attributed-QA.

As such, our contributions give some concrete answers to questions 1 and 2 above (How to measure attribution?, and How well do current state-of-the-art methods perform on attribution?), and give some hints as to how to address question 3 (How to build LLMs with attribution?).

Related Work

This section focuses on a few areas of related work.

Question answering has emerged as a key way to discover and demonstrate advances in LLMs. Reading comprehension asks a model to take as input a question and a passage which possibly contains an answer to the question, and to extract that answer. Since the seminal work of SQuAD Rajpurkar et al. (2016), there has been a proliferation of reading comprehension datasets developed to benchmark different machine capabilities that are important for QA Joshi et al. (2017); Choi et al. (2018); Reddy et al. (2019); Rodriguez et al. (2019).

The Natural Questions Kwiatkowski et al. (2019) effort provided a large reading comprehension dataset based on real information-seeking queries to the Googlehttps://www.google.com search engine, and has served more recently via Open-NQ Lee et al. (2019) as a benchmark for open-domain QA Voorhees and Tice (2000); Yang et al. (2015). In open-domain QA, a system receives only an input query and must return an answer based on a set corpus of evidence. Open-domain QA was first approached using retrieve-then-read pipelines, which use a trained retrieval engine to identify relevant passages, before performing reading comprehension over these to deduce an answer Chen et al. (2017). Both retrieval and reading comprehension have been actively investigated, e.g. using neural indexing Tay et al. (2022); Wang et al. (2022), dual encoders Ni et al. (2021) and few-shot prompting Chowdhery et al. (2022). Retrieve-then-read architectures are proposed as one class suitable for Attributed QA in Section 4. Concurrently, dense methods that jointly optimize for passage retrieval and answer prediction Lee et al. (2019); Karpukhin et al. (2020) have been successful, typically with less training signal than the pipeline approaches.

Roberts et al. (2020) shows that T5 Raffel et al. (2020) can perform a new task formulation, closed-book QA. Concretely, T5 can produce answers to questions without access to any corpus at inference time, instead producing answers based on its model parameters, tuned to “remember” information digested in pretraining. This result is tantalizing because it opens the possibility of more powerful question answering than so far realized in proposed task datasets. However, it requires us to fundamentally rethink how we approach question answering and its evaluation. We defer discussion of the advantages and disadvantages of the closed-book setting to the next section.

One result of the Natural Questions dataset is that we have a subset of examples implicitly gold-labeled for attribution. NQ produced examples of the form (x,a,c)(x,a,c), where xx is a question, aa is short answer, and cc is a long answer (typically a paragraph) selected by annotators as support for the answer aa. However, for a given question only a single long answer cc is annotated, so the set of attributed answers may be a small subset of those available on Wikipedia (the corpus also considered in our experiments). Extending from this, Petroni et al. (2021) describe the KILT benchmark for knowledge intensive tasks (including question answering), where gold-labeled “provenance” paragraphs are provided. Petroni et al. (2021) extends the NQ corpus’s coverage of provenance by using Amazon Turk annotators to mark additional paragraphs that support a given answer. The result is an increase from 1 provenance passage per (question, answer) pair to an average of 1.57 passages.

2 LLMs with Attribution

Existing work has explored whether attribution may be achieved using retrieval. Gao et al. (2022) proposes a two-stage technique where LLM-generated text is post-edited to be made attributable to web content retrieved in a first stage. Menick et al. (2022) propose a system, GopherCite, and perform a close set of evaluations to ours. In GopherCite, an LLM generates answers to questions using evidence retrieved by Google search for the given query as its input, along with paragraph-level supporting evidence that is evaluated using human raters. The guidelines are not specified in precise detail, though appear similar to AIS.

GopherCite is an important reference, but we see two limitations. First, only a single system, the hybrid Google search/LLM system, is evaluated. Second, the evaluation is limited. Of 307 questions presented to raters, only 115 questions were retained for evaluation, with the remaining 192 (62.5% of questions) being discarded due to raters skipping some items. Little is said about the basis on which raters skipped items, but this makes the 80% accuracy of the system hard to interpret.

There has been a recent flurry of activity in producing compelling proof-of-concept demos that generate seemingly factual responses in information-seeking dialog settings Nakano et al. (2021); Glaese et al. (2022); Thoppilan et al. (2022); Guu et al. (2020). These operate by incorporating a retrieval system, typically a commercial search engine such as Bing,https://www.bing.com/ into an LLM that then conditions its output on the retrieved content. The promise of these demos is a key motivation for this work: we perform a systematic study of different architecture decisions, using the principles in the AIS work Rashkin et al. (2021): recent demos most closely fit into the Retrieve-then-read class of architectures in Section 4, where other possible design choices are described, and each operationalize the concept of attributability differently.

Attributed Question Answering

This section defines the Attributed QA task, and gives discussion.

We assume a set C{\cal C}, which is a fixed set of units to which answers can be attributed. For example, C{\cal C} might be the set of all paragraphs in some corpus. More specifically, each c∈Cc\in{\cal C} is the ID for some unit; we use text(c)\hbox{\tt text}(c) to refer to the actual paragraph text for natural language datasets. The input to an attributed QA system gg is a question xx. The output from the system is a pair g(x)=(a,c)g(x)=(a,c), where aa is a text string, and cc is a member of C{\cal C}.

2 Evaluation

We consider two evaluation metrics for the Attributed QA task: first, human ratings that are the gold-standard, and second, automatic evaluation methods, which we show can be suitable in development settings. Section 5 gives analysis of the correlation between the two.

Given a triple (x,a,c)(x,a,c), we use the AIS evaluation definitions and guidelines Rashkin et al. (2021) to judge whether the answer to question xx is attributable to cc. Raters are asked to answer the following two questions, in the context of the question xx (where the system response is the answer aa, and the source document is cc):

Is all of the information relayed by the system response (a,c)(a,c) interpretable to you?

Is all of the information provided by the system response aa fully supported by the source document cc?

We define the rating of (x,a,c)(x,a,c) as “attributable” if the answer to both of these questions is “yes”.

Assume a set of test questions x1…xnx_{1}\ldots x_{n}, and a system gg to be evaluated. Define rir_{i} to be the (randomly chosen) pool of raters on the iith test example, and h(xi,g(xi),ri)h(x_{i},g(x_{i}),r_{i}) to be 11 if the majority of the annotators mark the system output g(xi)g(x_{i}) to be attributable, or . The test accuracy is then

That is, the test accuracy is simply the proportion of test examples where the majority of the raters judge the system’s output to be attributable.This is an estimate of E[h(X,g(X),R)]E\left[h(X,g(X),R)\right] where the expectation is taken over the random choice of example XX, and the random choice of raters RR.

Automatic Evaluation (AutoAIS).

In addition, we will make extensive use of an automatic measure, based on the NLI classifier of Honovich et al. (2022), AutoAIS Gao et al. (2022). See Section 5 for full details of the classifier. Taking AutoAIS(xi,g(xi))\hbox{\tt AutoAIS}(x_{i},g(x_{i})) to be the output of the NLI classifier (11 for attributable vs for non-attributable), we define

3 Discussion

Given this definition, we make the following remarks concerning the complexity of the task, motivation for the task, and the relationship to the more general attributed LLM problem.

We note that for many queries, the set of labels (a,c)(a,c) that form a correctly attributed answer to xx is likely to be large. This is due to two causes. First, for a given answer aa, there may be multiple paragraphs c∈Cc\in{\cal C} that support that answer. Second, for many queries, there may be a diverse set of answers that have some supporting paragraph in C{\cal C}. This diversity comes from several sources, for example: the same underlying answer being expressed by different strings; differing opinions about the answer to a question; differing answers under differing interpretations of a query. This makes evaluation of QA systems—whether or not attributed—challenging.

Remark 2: Motivation for Attribution.

Rashkin et al. (2021); Thoppilan et al. (2022); Menick et al. (2022) give extensive motivation for attribution in LLMs. We focus on a few key points here. First, attribution allows either a system developer or user to see the underlying source supporting an answer, and to assess aspects including trustworthiness and nuance. As such, attribution deliberately avoids the need for judgments of the “factuality” of claims, something that is challenging for all but the most simple questions, see Rashkin et al. (2021).

Second, attribution offers system developers a more streamlined human evaluation of answer quality. Consider instead QA definitions where a model simply outputs an answer string. Evaluation of new answer strings would require either:

Humans to use a search engine to attempt to find evidence for the answer. This places significant onus on raters and may be intractable;

The curation of one or more gold labels yiy_{i} for each test example xix_{i}, together with test error defined as 1/n∑i=1nL(g(xi),yi)1/n\sum_{i=1}^{n}L(g(x_{i}),y_{i}) for some similarity measure LL over answer strings. This is the closed-book QA setting.

Remark 3: Comparison to Closed-Book QA Evals.

The closed book QA setting has been valuable in developing LLMs and QA systems Roberts et al. (2020), and provides a highly effective and convenient measure of LLM performance. They do, however, have two significant drawbacks:

Closed-book QA evals do not require a system to provide attribution for its answers. In this sense closed-book evals measure performance on a task that is arguably incomplete or of limited utility to users and system designers.

Closed-book QA evals depend on gold-curated labels yiy_{i} for test examples, leading to significant difficulties for questions with diverse answers. In this sense there is a risk that closed-book evals significantly undercount performance (but without large-scale human evaluations such as those described in the current paper, it is impossible to estimate the scale of this problem).

Remark 4: Motivation for Human Ratings.

Throughout this paper we will take the final measure of system performance to be based on human ratings. A primary motivation for this is that given the size of the space of attributable (a,c)(a,c) pairs (Remark 1), it is unclear whether curating gold standard (a,c)(a,c) labels that cover enough of the output space is feasible for the task: or at least if attempts are made to do this, we will need to correlate with human evals to measure their effectiveness.

One side-effect of the large quantity of human labels gathered in this paper is that we may be in a much better position to develop high quality automatic evals for attributed QA. For example, we can measure the level of correlation between existing measures such as AutoAIS and human ratings. Or we can use the labels gathered to train automatic evals (see Sellam et al. (2020); Bulian et al. (2022)), potentially with a large set of reference (answer, attribution) pairs for each dev/test example.

Remark 5: Relationship of Attributed QA to Attributed LLMs.

Attributed QA is perhaps the simplest possible attributed LLM task, but it gets at the core task of attribution of “statements” or “propositions”: see Rashkin et al. (2021) for definitions and discussion. In short, the problem of attributing a question/answer pair (e.g., x=x= “when did the first dinosaurs live”, a=a= “230 million years ago”) is closely related to the problem of attributing statements made by an LLM, where a “statement” is some declarative sentence, for example “the first dinosaurs lived 230 million years ago”. Much of LLMs’ outputs in more complex information-seeking scenarios such as dialog and multi-faceted QA involve sequences of such statements (or, essentially equivalent, answers to questions), many of which require attribution. There are undoubtedly complexities in extending results for attributed QA to the full attributed LLM problem—for example deciding which statements need to be attributed, dealing with the move from question/answer pairs to more general statements, or dealing with complex statements that may involve attribution to multiple sources. But our working hypothesis is that progress on Attributed QA will extend naturally into more complex tasks.

Approaches to Attributed QA

We now describe the different systems investigated in this paper. At a high level, they fall into the following three architecture classes, and may be differentiated in terms of the type and quantity of supervision that is used. We will study a variety of different systems that fall into these three categories, carrying out ablations of key components.

Following approaches to open-domain QA, retrieve-then-read (RTR) models first perform retrieval of kk relevant passages based on the input question alone, where kk is a relatively small number. A second-stage model then takes P⊂kP\subset k retrieved passages, possibly reranked, as input to generate a short answer, and chooses one of A⊂kA\subset k retrieved passages as support for that answer.

In our experiments, we used BM25 Robertson and Zaragoza (2009) for sparse retrieval, GTR Ni et al. (2021) for dense retrieval, and Fusion-in-Decoding (FID, Izacard et al., 2022) for answer generation. FID may be trained with T⊂kT\subset k retrieved passages as input to answer generation, to reduce memory requirements. GTR may be used in the open-source version (PT-GTR in following tables), or further tuned for NQ (GTR).

Post-hoc retrieval.

In these systems, an LLM is first used to generate an answer to the input question, typically using few-shot prompting, without any use of retrieval. The question and answer are then concatenated to form a query to sparse or dense retrieval, again giving kk relevant passages. For k>1k>1, a final step selects the highest scoring passage containing the answer generated by the LLM as the attribution.

LLM-as-retriever.

In LLM-as-retriever models Tay et al. (2022); Wang et al. (2022), an LLM is used to generate both an answer and a pointer into the attribution corpus through some combination of prompting and fine-tuning but without any use of either sparse or dense retrieval. In this paper, we split attribution into a two-stage process, with the LLM first generating a webpage URL, from which a paragraph is then selected as support for the answer. A natural extension of this approach (not investigated in here) would be to generate a pointer to a paragraph rather than a URL.

2 Supervision

A second important axis in which systems can be differentiated concerns the type and quantity of supervision that is used. In NQ-64 systems, very limited supervision, in the form of 64 randomly chosen training examples from the Natural Questions (consisting of question/answer pairs) is used. In NQ-full systems, we assume access to the full NQ training set. In RTR pipelines, GTR retrieval and FID answer generation use NQ-full. When NQ-full is used to select exemplars in post-hoc pipelines, the 64 most similar examples to the target based on the BM25 score are used.

3 Best Systems

The next section of the paper describes experiments with a number of systems. We briefly highlight four particularly important ones (and two variants) that achieve the highest AIS score for their architecture.

GTR is first used for retrieval of top k=50k=50 passages from an input query, and the NQ-reranker is then used to rerank these passages. FID is trained with T=50T=50 passages but generates an answer based on the top P=1P=1 passage, which is returned as the attribution. (Note that this approach is very close to the approach of Izacard et al. (2022), with the added final passage selection step that allows evaluation for attribution.)

Best post-hoc retrieval system.

A prompted version of a 540B parameter PaLM produces an answer to the question. The prompts are 64 question/answer pairs from the Natural Questions training set, chosen based on BM25 similarity. GTR is then used for retrieval of an attribution, using the question concatenated with answer as the input query, by selecting the passage in the top kk = 50 which contains the PaLM-predicted answer string.

Best low-resource system.

A prompted version of a 540B parameter PaLM produces an answer to the question. Again, prompts are 64 question/answer pairs selected from the full Natural Questions training set. BM25 is used for post-hoc retrieval, again with the question/answer pair concatenated to form the query. We refer to this system as “very close to unsupervised” as it only requires 64 NQ examples, it does not require fine-tuning, and the underlying retrieval method does not require supervision or fine-tuning.

Best LLM-as-retriever system.

We explore the possibility of more end-to-end approaches to Attributed QA by fine-tuning a 540B parameter PaLM to generate an answer and Wikipedia URL, given a question input. The fine-tuning data used for this was questions generated by a question-generation model trained on SQuAD Rajpurkar et al. (2016) over decontextualized Choi et al. (2021) Wikipedia sentences, as well as the Natural Questions training set. The Wikipedia paragraph with the highest BM25 score is used as the attribution.

AutoAIS reranked variants.

To fairly assess how good RTR and PaLM post-hoc pipelines are at producing answers which could be attributed, we additionally experiment with system variants where AutoAIS is used as a reranker. These are identical to the above RTR and post-hoc systems, but instead use AutoAIS scores to select the attribution passages: the retrieved passage in the top k=50k=50 with highest AutoAIS score is selected as the attribution (and so is used to generate the answer in RTR). Since these variants use AutoAIS as a system component, we evaluate performance only on AIS (not AutoAIS). We encourage others who make such use of automatic evaluations to improve system quality to similarly distinguish between when they are being used as system components and when they are being used for evaluation.

Experiments

We now describe experiments on Attributed QA. We first give technical details, then present system results, before concluding with an analysis of the evaluation metrics.

We evaluate the short-answer seeking questions from the validation set of the Natural Questions Kwiatkowski et al. (2019), i.e. those that appear in OpenNQ Lee et al. (2019).

Attribution Corpus.

We use a snapshot of Wikipedia from 2021-10-13 to derive C{\cal C}, using Pyserinihttps://pypi.org/project/pyserini/ to extract paragraphs from each page.

2 Evaluation Metrics

We report three metrics for all experiments.

The gold-standard metric is the AIS measure assessed by human raters, as described in Section 3.1. Raters are trained using repeated annotations with feedback, until reaching high performance on the task and we take the majority vote from 5 raters. Given the cost of human rating, we evaluate on 1000 randomly-chosen questions and estimate standard errors using two-sided bootstrap re-samplinghttps://docs.scipy.org/doc/scipy/reference/generated/scipy.stats.bootstrap.html.

AutoAIS

AutoAIS formulates evaluation as a Natural Language Inference task that asks a model whether the question and answer are entailed by the provided attribution. We use a T5 Raffel et al. (2020) checkpoint with 11B parameters fine-tuned on a collection of NLI-related tasks (Williams et al., 2018; Bowman et al., 2015; Thorne et al., 2018; Zhang et al., 2019; Khot et al., 2018; Schuster et al., 2021). We score a given (premise, hypothesis) input by measuring the output probability when force-decoding the positive label, resulting in a score between 0 (no entailment) and 1 (entailment). We treat values ≥0.5\geq 0.5 as indicating valid (answer, attribution) predictions.

Exact Match (EM)

Finally, for comparison to prior work, we also report EMhttps://github.com/google-research/text-to-text-transfer-transformer/blob/2ce1574a0c2f5ed65a08e87cc38ad8ceb222b239/t5/evaluation/metrics.py#L154 for the answer string alone, ignoring attribution.

3 System Results

Table 1 shows results for the systems in each architecture class with the best AIS score, with AutoAIS Reranked variants. The most striking result is that the systems which perform best on AIS do not necessarily achieve the strongest EM accuracy (cf. Tables 2 and 3). This is discussed below in Section 5.5, where we find EM correlates only modestly with human judgment of AIS and has important limitations for Attributed QA evaluation. At the same time, we note that we did no special modeling to maximize EM score, such as instruction tuning Wei et al. (2021) or chain of thought prompting Wei et al. (2022), and that models tuned for greater EM may also achieve higher AIS scores.

Best RTR achieves the highest performance (p≪10−5,t=4.55p\ll 10^{-5},t=4.55, in comparison with the best non-RTR system), despite using LLMs with relatively small numbers of parameters (using T5 XL with 3B parameters, compared to PaLM with 540B). However, RTR approaches have the shortcoming that they require relatively large amounts of explicit supervision, for example in the form of NQ examples (an open question is whether RTR systems with much less supervision can be developed). They are also likely to be highly dependent on the accuracy of the retrieval step.

It is encouraging that Best Post-hoc achieves relatively high EM because it requires minimal amounts of supervision for answer generation (using prompting). However, these models generally require LLMs with large numbers of parametersFor example, Chowdhery et al. (2022) appendix H.1 reports NQ exact match results of 14.6/27.6/39.6% for 8/64/540 billion parameter models, showing significant performance increases with increasing scale. (presumably needed for memorization). Also, attribution poses a challenge in this setting; as noted above, on AIS the best RTR system is significantly better than the best post-hoc system, and this difference carries over to AutoAIS as well. However, since reranking is able to find good attribution passages, this result suggests that attribution is more difficult in a post-hoc setting than in RTR, and is a key area for future development.

Best Low Resource performs competitively with Best Post-hoc on AIS and AutoAIS despite using a sparse retrieval. This is promising for more complex information-seeking tasks where it is challenging to provide explicit supervision, and where LLMs have been shows to provide fluent output.

End-to-end models have the potential benefit of not requiring retrieval at all. That the performance of Best LLM-as-retriever is competitive with low-resource post-hoc attribution is promising, given that it is BM25 which is used to select a paragraph from the returned URL. However, they again require LLMs with large numbers of parameters.

4 Ablations

Ablation studies are presented for RTR systems in Table 2 and for post-hoc retrieval systems in Table 3. In the RTR models, the best dense-retrieval system (RTR-10) outperforms the best sparse-retrieval system (RTR-4) by 17 points AIS (p≪10−13,t=7.79)p\ll 10^{-13},t=7.79). Among the post-hoc systems, the dense retrievers also have the edge with the AIS difference between the best systems of each class (Post-6 vs Post-2) being statistically significant (p≪0.01,t=2.91p\ll 0.01,t=2.91).

For RTR systems, training FID with T=50T=50 examples seems essential for achieving top performance, though using P=50P=50 passages for answer generation is only useful if all A=50A=50 passages are considered for attribution also. Simply selecting the top retrieved passage (A=1A=1) as the attribution after training and generating an answer with 50 passages performs poorly (e.g., p≪10−7,t=5.60p\ll 10^{-7},t=5.60 for RTR-12 vs RTR-11 on AIS). That is, while the best performing architecture, RTR is resource intensive and it is unclear how to reduce this without hurting performance.

Across the post-hoc systems, selecting an attribution passage among the k=50k=50 that are retrieved seems better than simply using the top k=1k=1 (e.g., p≪.01,t=2.78p\ll.01,t=2.78 for Post-5 vs Post-6 but p=0.04,t=2.01p=0.04,t=2.01 for Post-1 vs Post-2). The interesting trend here is in the impact of expanding the pool of NQ examples for exemplar selection. Using NQ-full gives a 10 point boost to Exact Match for both BM25 and GTR systems but the impact on AIS is much smaller.

Taken together, these results show that existing state-of-the-art methods are suitable for Attributed QA, though there is still headroom to improve, especially in the post-hoc attribution of LLM-generated answers. As to how to design systems, we have discussed how it depends on many factors which should be carefully considered.

5 Correlation between AIS and EM/AutoAIS

We now focus on the question of how to best measure attribution given our observations so far. To do this, we estimate the correlation between system scores on (human) AIS, and EM and AutoAIS in turn, by calculating the Pearson coefficient between the two sets of scores (i.e. between AIS and EM scores, and between AIS and AutoAIS scores).

We saw above that best AIS performance did not necessarily go hand-in-hand with best EM accuracy. Consistent with this, the Pearson correlation coefficient between the system EM and AIS scores is modest, at 0.45 (see Figure 2). Manual analysis of the disagreements revealed multiple factors to be involved, including answers with inexact string matches to the NQ reference answer, stale reference answers, and questions with more than one valid answer able to be retrieved (see Table 4). Overall, we suggest that our results point to the limitation of reference answer corpora and string matching evaluation for future research.

AutoAIS

On the other hand, correlation between system AIS and AutoAIS scores is remarkably strong, with a Pearson coefficient of 0.96 (Figure 3). This suggests that AutoAIS is fit-for-purpose as a development metric at the aggregate level (provided it is not used as a system component).

To get a deeper understanding of the correlation, we followed up with an instance-level correlation study where the data series are the per-question ratings for a given system. Correlation was much lower and more variable here. Therefore, we recommend care should be taken against reading individual AutoAIS scores too closely.

The reranked variants are outliers to this strong correlation, with attributions selected by AutoAIS scoring lower on human evaluation than would be expected based on a linear fit. This is consistent with instance-level AutoAIS (which was used in reranking) being noisier than system-level AutoAIS: the passage with the best AutoAIS is not necessarily the one preferred by humans.

Future Directions

We see many exciting areas for future work.

While retrieve-then-read systems achieve strong performance, this class typically requires a large amount of data to train and can be resource intensive. We are excited by the possibility of post-hoc attribution of LLM-generated answers and end-to-end modeling for Attributed QA. Future directions to improve performance in these settings includes studying the challenge of retrieval for post-hoc attribution, and devising training signals for end-to-end modeling. One possible, albeit noisy, source for the latter is AutoAIS, which we observed correlated well at the system-level with human judgments of AIS. We also noted the promise of instruction tuning and chain of thought prompting for improving the quality of LLM-generated answers.

Evaluation.

We observed that AutoAIS was fit-for-purpose as a development metric, but had shortcomings including only moderate correlation with human ratings at the instance-level. There are at least two possible ways to use the human rating data collected in this paper to improve from this. First, the data could form a cache used to score system predictions which have been observed previously. In this way, the data could be seen as an extension of KILT Petroni et al. (2021), curating a range of attributed answers that do not require further verification. A softer approach could apply prior work Sellam et al. (2020); Rei et al. (2020) and use the data to learn an improved automatic evaluation metric for attribution. We note that the latter is additive with using AutoAIS as a noisy training signal for end-to-end learning.

Tasks.

We have presented an in-depth study on the Natural Questions to demonstrate the promise of an attribution task with automatic and human verification. However, our best system requires use of the full NQ training set. We would like to understand how general-purpose this approach is, and whether systems that make less use of direct supervision transfer to new settings better. Therefore, we see future work in evaluating on different datasets (esp. Joshi et al., 2017), perhaps with multilingual Clark et al. (2020) or multimodal Antol et al. (2015) attribution. We are excited by the challenge of attributing generated text more generally, perhaps in long-form QA Stelmakh et al. (2022).

Conclusion

We establish a research agenda to develop attributed large language models. We believe that attribution will be crucial for technologies based on LLMs in information-seeking settings. To understand how to make progress in this area, we define and study a new task, Attributed QA, which bases evaluation on the AIS principles and benchmarks architecture designs using a range of state-of-the-art components to build systems. We consider human rating to be the gold standard for system evaluation, but find that AutoAIS correlates well with human judgment at the system level, offering promise as a development metric where human rating is infeasible, or even as a noisy training signal. Retrieve-then-read approaches achieve the strongest performance on our evaluation, but require full use of a traditional training set. Post-hoc attribution appears to be a viable architecture for future work, but remains challenging.

Ethical Considerations

The main ethical consideration of this work concerns "factuality." As in Rashkin et al. (2021), we observe that it is incredibly challenging to judge whether any but the simplest claim is factual. Instead, for most questions, there will be multiple valid answers that are distinguished by nuances that can be subtle. Therefore, we believe attribution will be crucial in most information-seeking scenarios and explore what it means for an LLM to be able to attribute text it generates. In this way, users can inspect sources to make their own judgment of trustworthiness and answer scope. It is an interesting research question not studied here, how to identify issues like factual inaccuracies and biases in web sources.

We also consider the issue that Attributed QA is only explored in English using, for the most part, resource-intensive approaches that may not be accessible to many. To encourage future work that expands from here, the AIS principles are publicly available Rashkin et al. (2021) and we have released all system outputs and their ratings. We are excited by the promise of low-resource and end-to-end solutions to meet the diverse challenge of attribution in language modeling.

Contributions

Bernd Bohnet, Vinh Q. Tran, and Pat Verga lead the technical work for this paper, including implementing models, running experiments, analyzing results and making improvements. Kellie Webster acted as TL.

Livio Baldini Soares built the core infrastructure for BM25 retrieval, Ji Ma and Jianmo Ni contributed the dense retrieval and reranking pipelines, and Kai Hui helped with PaLM usage. Daniel Andor and Kuzman Ganchev ran components that enabled the LLM-as-retriever model.

Bernd Bohnet managed the human rating collection and its data pipeline. Roee Aharoni and Jonathan Herzig trained the NLI models for automatic evaluation and Roee open sourced the model on huggingface for public use. Massimiliano Ciaramita conceived the AutoAIS prompt and Lierni Sestorain Saralegui explored several variants. Tom Kwiatkowski, Livio Baldini Soares, and Daniel Andor built the Attribution corpus and helped with datasets. Kellie Webster implemented the standard automatic evaluation script around this work.

Michael Collins and Kellie Webster were the primary writers of the paper. William W. Cohen helped with the related work, Jacob Eisenstein contributed the statistical analysis in the results section. William W. Cohen, Michael Collins, Dipanjan Das, Don Metzler, Slav Petrov, and Kellie Webster developed the direction for this work and contributed significant feedback during paper writing.

Acknowledgements

We would like to thank our many colleagues whose insightful discussion shaped this work, including Fernando Pereira, Ankur Parikh, Jon Clark, Marc Najork, and Vitaly Nikolaev. The human rating process was managed by Muqthar Mohammad and Isabel Kraus-Liang, who worked diligently to produce incredible results. Kathy Meier-Hellstern, Suneet Dhingra, and teams provided invaluable support.

References

Appendix A Examples

We identify three classes of interesting examples that demonstrated the value of AIS over EM.