A Critical Evaluation of Evaluations for Long-form Question Answering

Fangyuan Xu, Yixiao Song, Mohit Iyyer, Eunsol Choi

Introduction

Long-form question answering (Fan et al., 2019; Krishna et al., 2021; Nakano et al., 2021; Su et al., 2022, henceforth LFQA), an emerging research area within QA, requires systems to generate long and complex answers to questions by leveraging large language models and evidence document retrievers. While remarkable strides have been made in LFQA model development, the current state of LFQA evaluation is dire: most prior papers use a combination of crowdsourced human annotations and simple string-matching metrics (e.g., ROUGE). We present the first study of the evaluation of long-form answers, exploring both human and automatic evaluation protocols to better understand how we should evaluate LFQA moving forward.

Human evaluation: In most prior human LFQA evaluations (Krishna et al., 2021; Nakano et al., 2021), crowd annotators are given a question, two candidate answers, and (optionally) evidence documents, and they are asked to identify the better answer. However, crowdworkers do not necessarily have the expertise or background knowledge to reliably judge properties such as factuality (Gillick and Liu, 2010; Iskender et al., 2020). Thus, we hire domain experts in seven different fields (e.g., biology, economics) to perform the same answer preference task and additionally provide detailed justifications as to why they chose a particular answer. Analyzing their justifications reveals that experts consider properties such as completeness and factuality to be more decisive than surface-level aspects (e.g., conciseness and level of detail) on which crowdworkers tend to fixate. Additionally, even experts often disagree with each other about which answer is better; this disagreement stems from valuing fine-grained answer properties differently.

Automatic evaluation: As human evaluation is slow and expensive, developing a reliable automatic LFQA evaluation metric is crucial for speeding up model development. While ROUGE Lin (2004) has been shown to be misleading for LFQA Krishna et al. (2021); Wang et al. (2022), do any other existing text generation metrics correlate to human judgments of answer quality? Can we train a metric to mimic human preference judgments? To answer these questions, we curate a suite of 12 automatic metrics and measure how they correlate to human judgments of both “overall quality” and two fine-grained aspects (coherence and faithfulness). None of these metrics reliably matches human judgments of overall answer quality. However, automatic metrics such as QAFactEval Fabbri et al. (2022) and RankGen Krishna et al. (2022) show potential at modeling fine-grained aspects of LFQA answers, which can spur research on a new generation of automatic LFQA metrics.

Overall, we provide the first thorough study of LFQA evaluation and shed light on the components of good long-form answers. As part of our exploration, we collected and will release a small-scale dataset of expert evaluation of long-form answers (260 ratings and justifications over 140 answer pairs). We conclude by providing recommendations for the future of human and automatic LFQA evaluation, encouraging the community to hire expert evaluators and move from poorly-defined judgments of “overall preference” to a multi-faceted evaluation modeling attributes such as answer completeness, factuality, and ease of understanding.

Background and related work

We begin by reviewing the evaluation protocols used by prior work in LFQA, which has centered around a dataset scraped from the “Explain Like I’m Five” subreddit (Fan et al., 2019, ELI5).https://www.reddit.com/r/explainlikeimfive We include brief review of evaluation in other text generation tasks in Appendix A.1.

Early work on LFQA Fan et al. (2019) uses ROUGE (Lin, 2004) to measure the similarity of human reference answers to model-generated answers. Krishna et al. (2021) find that ROUGE is not a meaningful metric due to the open-ended nature of long-form answers, but they do not examine other automatic metrics. Given the difficulty of evaluation, recent works re-scoped the task to allow more reliable evaluation: Wang et al. (2022) focus on exemplification in long-form answers by treating this sub-task as a retrieval problem, while Stelmakh et al. (2022) aim to evaluate long form answers limited to ambiguous factoid questions that cover the different disambiguated questions and their corresponding answers. However, these evaluation protocols cannot be easily adapted to the general LFQA task: the metric in Stelmakh et al. (2022), for example, requires a list of disambiguated questions and their answers, which is not available for many questions.

We summarize the human evaluation studies conducted by two previous studies, Hurdles Krishna et al. (2021) and WebGPT Nakano et al. (2021). Both works evaluate via A/B testing (i.e., choose which of two candidate answers is better), and they collected judgments of overall answer quality, factuality, and coherence. While both works recruited non-expert annotators and collect only one-way annotations, WebGPT’s evaluation allows annotators to look at a set of evidence documents when judging the answer, and they also collect optional free-form justifications from the annotators to justify their choice. While fine-grained aspects such as coherence Goyal et al. (2022); Jiang et al. (2022) and factuality Goyal and Durrett (2020); Laban et al. (2022) have been studied before for other tasks such as summarization, ours is among the first works to study LFQA-centric properties such as completeness or ease of understanding.

How do domain experts evaluate long-form answers?

Prior LFQA human evaluations use non-expert crowdworkers to evaluate highly domain-specific answers, either with no access to external information (Krishna et al., 2021) or access to only model-retrieved evidence documents (Nakano et al., 2021). Both settings are problematic: non-experts cannot be relied on to judge the correctness of answers in isolation, and they also cannot be expected to thoroughly comprehend evidence documents and judge their validity or relevance to the answer (Gao et al., 2022). While Nakano et al. (2021) solicit optional free-form justifications from their workers to explain their preference judgments, it remains unclear how well these workers can judge correctness in fields that are not their expertise. Our first contribution is to hire domain experts in seven fields (see Table 2) and have them evaluate both human-written and model-generated answers via A/B judgments as well as paragraph-length free-form justifications. An analysis of the expert annotations reveals a complex and subjective interplay between many different fine-grained aspects of LFQA answers (e.g., completeness, factuality) that pose challenges for future LFQA evaluation.

We recruit domain experts on the freelancing platform Upwork for seven domains shown in Table 2. Each expert has earned at least a bachelor’s degree in the target domain and has expertise performing tasks in that domain (e.g., summarizing scientific articles or being a teacher of the domain). As shown in Table 2, we hire 1-3 experts per domain. Given a question and two candidate answers, the experts were asked to choose which of the answers is better (overall preference), indicate whether the decision was difficult to make (e.g., because both answers were of similar quality), and lastly to justify their choice in a free-form paragraph. The evaluation tasks are hosted on Label Studio.Figure 4 contains a screenshot of our annotation interface. https://labelstud.io/ The experts reported that they spent 15 to 30 minutes per question, which shows the demanding nature of the annotation task. We accordingly paid $3.253.25 per question, which resulted in a total cost of $845 to collect 260 expert judgements.We explain in Appendix A.2 why the numbers of experts in each domain differ.

Following prior work, we conduct A/B preference testing on two answers to the same question. We include two settings: (1) H/M: comparing a model-generated answer with a highly-upvoted human-written answer, and (2) H/H: comparing a highly-upvoted human-written answer to an answer with fewer upvotes (where upvotes are a noisy proxy to answer quality).Reddit users give upvotes to content to show their support or approval of the content. The first setting is intended to identify common classes of errors made by state-of-the-art LFQA systems, while the second setting is more of a sanity check exploring whether low-effort human answers make similar errors to models.

We chose GPT-3 text-davinci-002 model (175B) Brown et al. (2020b) as the LFQA model to evaluate. A small-scale qualitative analysis found that zero-shot GPT-3 possesses more advanced LFQA capabilities than fine-tuned LFQA systems built on smaller language models. Since this model may have already seen the entire ELI5 dataset released by Fan et al. (2019) during its pretraining, we scrape more recent questions from the r/explainlikeimfive and r/AskHistorians subreddits posted between July to December 2021.text-davinci-002 was trained on data up to June 2021. Question askers on the ELI5 subreddit often categorize their questions into domains via the flair label, which enables us to perform a domain-specific analysis.The details on domain identification are in Appendix A.2. We randomly sample 20 questions per domain except for the history domain, which has 15 questions in the H/M setting and 5 in H/H. This discrepancy is due to the difficulty of finding history questions with a moderate answer length. As shown in Figure 1 and Table 5, human-written answers to history questions are much longer than the answers in the other domains, even after careful screening.

To obtain model-generated answers, we prompt the model in a zero-shot manner with the following prompt: “Generate a long answer to the following question with examples and references when necessary.” For decoding, we used the default decoding setup in the API (i.e., top p=1p=1 and temperature=0.7=0.7).

2 Quantitative results

As shown in Table 2, experts surprisingly display a slight preference (61.8%) for model-generated answers from GPT-3 compared to human answers; as a sanity check, they exhibit preference (62.4%) for highly-upvoted human answers over those with fewer upvotes. The preference of our annotators for model-generated answers is corroborated by similar findings for summarization by Liu et al. (2022), who show that GPT-3 generated summaries score higher than reference summaries.

Comparing different domains, we observe that model-generated answers are strongly preferred in economics (90%) and law (also 90%), while human answers are preferred in the history domain (75.6%). To understand the divergence in preferences for different domains, we report the answer length distribution of both answer types in the H/M setting in our expert-annotated dataset in Figure 1. The model’s struggles in history domain are likely because this domain contains the longest and most complex questions as well as human answers (averaging 356 words long in the H/M setting) out of all domains. Table 5 in the appendix report the length of questions, model-generated, and human-written answers of the whole expert-annotated dataset.

We report Fleiss’ κ\kappa Fleiss (1971); Landis and Koch (1977); Fleiss et al. (2013) as a measure of agreement in Table 2. Our expert A/B testers achieved fair agreement in economics, moderate agreement in biology and physics, and a substantial agreement in history. We observe that agreement increases when comparing a high and low-upvoted human answer together, as opposed to comparing model-generated answers with human answers. We emphasize that disagreement is not a failure of one of the experts to properly evaluate the answers. In fact, disagreement within experts highlights the challenges (and futility) of judging “overall answer quality” in this way. There are many salient properties of long-form answers, which we discuss next, and deciding how to value each property when coming up with an overall preference is highly subjective (see Appendix Table 8 for several examples).

3 What makes one answer better than another?

To better understand the various components of a good long-form answer, we perform an analysis on the free-form justifications collected from both our expert annotators as well as WebGPT crowd annotators from Nakano et al. (2021). WebGPT allowed optional justifications, and many of them are not very long or detailed. Our justification is about three times longer on average (statistics can be found in Table 6 in the Appendix). Our analysis focuses on the model-generated vs. human-written answer setting, where the model is either zero-shot GPT-3 (our work) or the 175B WebGPT model. Concretely, we analyze 50 randomly sampled justifications from each population. Our analysis is limited in that these two comparisons do not consider the same set of questions. We identify and code nine fine-grained aspects that are mentioned in them, and mark whether these aspects are decisive factors for making the preference judgment. The results are summarized in Figure 2, and we highlight takeaways below.

Perhaps unsurprisingly, our experts mention factuality in their justifications almost twice as frequently as crowdworkers (36 to 20), and it is the most common aspect referenced by experts. As an example, in the first row of Table 1, the expert accurately points out incorrect information in Answer A about blood thinners breaking up clots. Since WebGPT annotators lack domain expertise, they generally judge factuality by checking if a statement is supported in evidence documents, which gives them only limited coverage over the full answer.

We observe that experts mention completeness as a decisive criteria twice as often than WebGPT annotators (12 vs. 6). Completeness refers to whether the answer adequately addresses all aspects of the question or provides all necessary information to clarify the question. Judging completeness requires deeper domain expertise than a handful of retrieved articles offer. As an example, in the second row of Table 1, the expert states that Answer B mentions only one reason why people go bald (hormonal), while Answer A mentions hormonal and environmental factors and is thus superior.The expert further points out that both answers miss a third major cause of baldness: genetics.

Both experts and crowdworkers mention easiness to follow as a decisive criterion at the same frequency; in fact, this is the most decisive aspect for both populations. One of the main goals of LFQA is to convey the answer of a question to a non-expert; as such, it makes sense that this property is so critical. We emphasize that this has never been evaluated in prior LFQA research and encourage future work to embrace it as a major component.

WebGPT annotators are far more likely to mark conciseness and specificity as decisive factors for their preferences than experts. They prefer shorter to-the-point answers, despite the fact that such answers might be incomplete, and they also prefer answers that include specific details instead of generalities. We note that these properties are much more feasible to judge for crowdworkers than factuality and completeness, which is likely a reason why they are mentioned so frequently (Table 10 in the appendix for examples).

3.1 Do models understand justifications of human preferences?

Our manual analysis of the justifications shows that experts consider a wide range of aspects when forming their decision. Detailed justifications of generated answers are useful in understanding why an answer was preferred, but they are costly to obtain. Generating these justifications automatically and evaluating them is outside the scope of this paper. Instead, we perform a simpler evaluation via a proxy task: given a justification with masked references to both candidate answers, can a model disambiguate the missing references? An example of the task is below:

Input: Question: qq Answer A: a1a_{1} Answer B: a2a_{2} Comment: Both answers are coherent, but Answer <extra_id_0\tt{extra\_id\_0}> is completely irrelevant to the question since it is about a bionic ear instead of a person learning speech when they get a hearing implant. Answer <extra_id_1\tt{extra\_id\_1}> is relevant and a complete, concise answer. Expected Output: <extra_id_0\tt{extra\_id\_0}> B <extra_id_1\tt{extra\_id\_1}> A

We experiment with pretrained T5 checkpoints Raffel et al. (2020) of different sizes (220M, 770M, 3B, and 11B parameters) on our task zero-shot.We experimented with two-shot prompting with GPT-3 but observed worse results compared to the outputs from T5-3B and T5-11B, potentially because the task resembles the pretraining setup of T5. For each (question qq, answer pairs (a1a_{1}, a2a_{2}), justification jj), we construct three types of inputs: Original: The original justification jj with (q,a1,a2)(q,a_{1},a_{2}), Flipped: The original justification jj with flipped answer identity (q,a2,a1)(q,a_{2},a_{1}), Random: jj with randomly paired q′,a1′,a2′q^{\prime},a_{1}^{\prime},a_{2}^{\prime}, as a baseline. We evaluate using token-level exact match, which gives the model credit only when its output exactly matches that of the target. We expect better than random performance on Original and worse than random performance on Flipped if the model comprehends the justifications.

Results are shown in Table 3. We see that T5-3B an T5-11B are able to comprehend the justifications, as they show different results for original and perturbed comments. This suggests adapting LMs for multi-faceted automatic evaluations of long-form answers is promising. Preprocessing details on this study are described in Appendix A.2.1

Do automatic metrics correlate with human judgments?

The experiments in the previous section establish that LFQA is very difficult for humans to converge on in terms of an “overall” score, as even domain experts disagree with each other when choosing a “better” LFQA answer. Furthermore, several properties of these answers are important to evaluate, including factuality, relevance, and coherence, among others. Do existing automatic text generation metrics correlate with human judgments of these fine-grained aspects, or“overall” answer preference? We now explore this question with a wide range of text generation evaluation metrics.

We experiment with existing text generation metrics and metrics that we train directly on the human preference judgments.

Prior work used existing text generation metrics (e.g., ROUGE) to evaluate LFQA. The metrics were initially designed for other text generation tasks (e.g., translation or summarization), and their usage has not been validated for LFQA.

Many generation metrics assume access to human-written references (in our case, gold answers), which are used to compute similarity scores to model-generated text. Of these, we evaluate ROUGE Lin (2004), which is the only reference-based evaluation metrics employed by prior work for LFQA, as well as BERTScore Zhang et al. (2019) and BLEURT Sellam et al. (2020), which leverage pretrained language models and have shown to be effective in evaluating many generation tasks (Kasai et al., 2022). A major limitation of reference-based metrics for LFQA is the huge space of valid output answers for any given question, which has been noted in prior work (Wang et al., 2022).

Some aspects, such as fluency and coherence, can be determined by looking at just the answers alone. Thus, we also examine a set of answer-only automatic metrics: (1) Self-BLEU Zhu et al. (2018), which measures the diversity of generated text (higher scores mean lower diversity) and has been previously used in open-ended generation Holtzman et al. (2019); and (2) GPT-2 perplexity, which prior work on constrained generation Zhang et al. (2020); Qin et al. (2022) has used to evaluate fluency.

Good answers should be relevant to the question asked, so we can model p(q∣a)p(q|a) to rank answers using the following methods: (1) Zero-shot question likelihood, which uses the instruction-tuned T0 model Sanh et al. (2022) to calculate the likelihood of the question given the long-form answer; (2) BARTScore (Yuan et al., 2021), which is an encoder-decoder model fine-tuned on text summarization; and (3) RankGen Krishna et al. (2022), which is an encoder model trained contrastively to score model-generated sequences (in our case, answers) given a prefix (the question).

Arguably the most challenging aspect of LFQA evaluation is to measure the correctness of the answer. While there are no existing factuality metrics for LFQA, the task is related to faithfulness in summarization. Metrics for faithfulness assume access to a set of evidence documents and evaluate whether a text is supported by the evidence Kryscinski et al. (2020); Goyal and Durrett (2020); Barrantes et al. (2020); Laban et al. (2022). We experiment with the QAFactEval metric Fabbri et al. (2022), which evaluates faithfulness by comparing answers from the summary (in our case, the answer) and the evidence document (retrievals from the WebGPT LFQA system).

1.2 Trained LFQA metrics

The metrics discussed so far are not trained on long-form answers. We now shift to training an LFQA evaluation metric directly on human-annotated preference judgments of pairs of long-form answers. Prior work from OpenAI Nakano et al. (2021) experimented with learning an evaluation metric by fine-tuning WebGPT to rank pairs of answers. As this model is not publicly available, we fine-tune a smaller-scale pretrained language model (176M Longformer-Base model) and rely on OpenAI’s API to fine-tune bigger pretrained language model (6B GPT3 text-curie-001 model.To the best of our knowledge, OpenAI has not clarified the exact size of each of the models in the API. We use this estimation:https://blog.eleuther.ai/gpt3-model-sizes/.) Details of fine-tuning setup are in Appendix A.4.1.

We use comparison data collected by Nakano et al. (2021) for fine-tuning, which contains 17,598 preference annotations. We remove ties and randomly split the data into train, validation and test sets with a 70%, 15%, 15% ratio. More details are provided in Appendix Table 12.

Our learned metric f takes in question qq, answer aa, and optionally evidence documents dd to produce a scalar score. We encode [q, a] and [a, d] separately with an encoder model and concatenate respective [CLS] representation then pass it to a linear layer to obtain a scalar score ss. As our input text is relatively long, we fine-tune a Longformer encoder Beltagy et al. (2020).

Following Nakano et al. (2021), we train the model with cross-entropy loss such that the scores produced by ff rank a pair of answers (a1a_{1},a2a_{2}) in the same order as the human preference. We estimate the likelihood that a1a_{1} is preferred over a2a_{2} as exp(s1)exp(s1)+exp(s2)\frac{exp(s_{1})}{exp(s_{1})+exp(s_{2})} where s1=f(q,a1),s2=f(q,a2)s_{1}=f(q,a_{1}),s_{2}=f(q,a_{2}). Given a set of answer pairs with gold preference p^\hat{p}, the loss is,

To leverage the advanced capabilities of larger-scale language models, we use OpenAI API to finetune GPT-3 text-curie-001 with the same comparison data split we used for the Longformer. Given a prompt consisting of question qq, answer a1a_{1} and answer a2a_{2}, the model is fine-tuned to output the label Answer1 or Answer2. This metric takes a pair of answers as input and outputs a preference, unlike the Longformer model which produces a score given a single answer.

2 Evaluating automatic metrics

Each evaluation example consists of {(q,a1,a2,p^)}\{(q,a_{1},a_{2},\hat{p})\}, where qq is question, a pair of long-form answers a1a_{1} and a2a_{2}, and p^\hat{p} ∈\in {a1a_{1}, a2a_{2}} denotes the human preference of choosing answer a1a_{1} or a2a_{2}. We report the accuracy of the metric preference pip_{i} against the gold human preference pi^\hat{p_{i}}. We omit the evidence documents d1,d2d_{1},d_{2} here for simplicity, but QAFactEval and longformer (D) metric take the evidence documents as additional input.

We compile human evaluations from previous studies Krishna et al. (2021); Nakano et al. (2021) and our expert annotations from Section 3. See appendix A.3 for descriptions of the models evaluated in these datasets as well as data statistics on the answers. Both prior studies present large-scale preference judgments of overall answer quality and smaller-scale judgments for two targeted aspects, coherence and factuality. In total, we look at 3,478 comparisons on overall answer quality, 854 comparisons on coherence, and 469 comparisons on factuality. As shown by our analysis of expert annotations (Section 3), annotators can frequently disagree with each other.

3 Results

Table 4 reports the accuracy of each metric at imitating human preference data. We report three baselines: Random, which randomly chooses one of the answers; Always Human, which prefers the human-written answer when available; and Length, which prefers the longer answer.The Length baseline is inspired by prior findings in summarization Sun et al. (2019); Liu et al. (2022) that length has a non-trivial impact in human preferences.

All metrics exhibit relatively low accuracies, falling substantially below estimated human agreement. None of the metrics are robust across different types of input answer pairs. For instance, pretrained reference-based metrics such as BERTScore and BLEURT have low accuracy on Hurdles human vs. model data, which adds further evidence to the issues with ROUGE noted by Krishna et al. (2021). Supervised metrics (Longformer and GPT-3) also struggle in this setting, despite outperforming all other metrics on overall rating in the other three data settings. While trained to imitate only overall rating, they achieve relatively strong accuracies on fine-grained ratings too, suggesting that they are correlated.

We observe spurious correlations with length for long-form answer evaluation. Choosing the longer answer achieves higher accuracy than all unsupervised metrics for the WebGPT model vs. model comparison; the best performance on factuality for Hurdles human vs. model answer; and the second-highest accuracy on our expert data. On the other hand, when comparing WebGPT human vs. model answers, choosing a shorter answer would have been more beneficial for coherence evaluation (62% of the time).The “strong” performance of the length baseline displays the brittleness of all existing automatic metrics for LFQA.

It is more feasible to model fine-grained answer aspects than overall answer quality. The QAFactEval metric, designed for factuality, does indeed outperform all other metrics on factuality. However, the metric is limited in that it requires a set of input evidence documents, which may not always be available or reliable. For coherence, simpler metrics such as self-BLEU perform competitively, and we also find that our upper bound of always choosing the human answer performs strongly on coherence, suggesting that models struggle to generate coherent long-form answers.

Correlation of Automatic Metrics Given pairs of long-form answers of the comparison data, we measure how frequently two automatic metrics prefer the same answer (Figure 3). We see a positive correlation among reference-based metrics (e.g., rouge and bertscore gives the same ranking for 63% of the pairs), as well as the (question, answer) metrics (e.g. qg likelihood and bartscore).

Conclusion & Future Work

Our study provides a unified evaluation benchmark for long-form answers, including new annotations from domain experts. We present a new set of expert LFQA evaluations along with detailed justifications, and we also compile existing human annotations across different properties (overall preference, factuality, coherence) to facilitate future development of automatic LFQA metrics.

Evaluation of long-form answers is a multi-faceted problem and thus should be more targeted. Our expert justifications suggest that many aspects are considered when deciding which answer is better, some of which may be at odds with others (e.g. completeness vs. conciseness). This suggests that computing an “overall” score for answer quality is not meaningful, which is further supported by the limitations of metrics trained directly from overall preference judgments. Future work should look deeper into modelling frequent aspects mentioned by expert annotators, such as completeness and ease of understanding, perhaps by taking inspiration from evaluation methods that explicitly localize and categorize errors Freitag et al. (2021); Goyal et al. (2022).

Limitations

We study a limited scope of long-form answers. The questions are either drawn from search queries or from community forums. In the real world, we will encounter many more diverse forms of long form question answering, such as answering questions in education or commercial settings. We only cover the English language, and thus our questions are topically limited to English-speaking culture.

Our evaluation of long-form answers is stationary. Annotators are provided a pre-generated output from the model without being able to interact with the model over multiple rounds. A more interactive evaluation Lee et al. (2022) of models is a great direction for future work.

Ethics Statement

The expert annotation data collection protocol has been determined to be exempt from review by an IRB board. All data collected will be made publicly available under the MIT license.

The data collection process did not require any information that can be used to uniquely identify individual workers. We examined the annotation data to make sure no such information or offensive content is present in questions or answers.

Acknowledgements

MI and YS were partially supported by awards IIS-1955567 and IIS-2046248 from the National Science Foundation (NSF). FX is supported by a fellowship from UT Austin. We thank the WebGPT team, especially Jacob Hilton, for sharing their human evaluation data with us. We thank the expert annotators for participating in our human evaluation. We thank Jessy Li and members of the UT Austin NLP community for helpful discussion to improve the paper. Lastly, we thank the reviewers and meta reviewer of ACL community for helpful comments and feedback on the paper.

References

Appendix A Appendix

Human and automatic evaluation for text generation is an active research area. We provide a brief overview here and direct the readers to recent surveys for more discussion Celikyilmaz et al. (2020); Gehrmann et al. (2022). Many tasks such as machine translation and summarization primarily rely on reference-based evaluation, with metrics such as BLEU Papineni et al. (2002), ROUGE Lin (2004) and BERTScore Zhang et al. (2019). These metrics aim to measure similarities between generated text and reference text. For open-ended generation problems such as story generation, comparing the generated text with a single reference is not meaningful. Reference-based metrics which instead measure the distributional similarity of model-generated and human-written texts have been proposed Pillutla et al. (2021). There has also been work on reference-less metrics, which mostly measure a specific aspect of text. For instance, factuality metrics for summarization Goyal and Durrett (2020); Kryscinski et al. (2020); Barrantes et al. (2020); Laban et al. (2022) capture the relationship between source document and summary, without the need of a reference summary. Another line of work proposes automatic metrics which learn to emulate human judgements of generated text, using either gold human preference or synthetically generated data Sellam et al. (2020); Zhong et al. (2022); Zhang et al. (2022).

A.2 Expert Annotation

Four domains (biology, physics, chemistry, and economics) are marked in the ELI5 posts (i.e., flairs), and two (tech/cs and law) are identified by using a dense passage retrieval Karpukhin et al. (2020) and KMeans from scikit-learn Pedregosa et al. (2011). Specifically, we use DPR to encode question of all posts whose flair is marked as others. Then, we run KMeans to find two big groups of questions whose domains can be reliably marked as tech/cs and law.

Experts are hired based on their academic background and English proficiency. No other demographic and geographic restrictions were applied. For each question domain, we aimed to hire three domain experts who have at least a bachelor’s degree in the domain through a paid pilot study. Thirty-five potential experts participated in a paid pilot study with 5 question-answer pairs. We paid $33 per question-answer set. At the end, only 13 experts met the qualification requirements and were willing to continue because the task required substantive expertise as well as time and attention commitment.

A.2.1 Justification Analysis

Data statistics of explanations collected are in Table 6. Examples of explanation and extracted aspects in our manual analysis can be found in Table 7.

To construct the masked comments, we first preprocess the justifications such that all mentions of the answer entity is prepended with the word “Answer” (i.e. replacing “Option A”, “A” with “Answer A”). We then mask out any mentions of “A” and “B” in the comment. We remove comments that do not contain answer entities after preprocessing, resulting in 259 (out of 260) expert comments and 292 (out of 305) WebGPT comments.

A.3 Previously Collected Human Evaluation Data

Dataset statistics is shown in Table 9. We group the comparisons by whether they are (model-generated answers v.s. human-written answers) or (model-generated answers v.s. model-generated answers), and present overall statistics. The model-generated answers include four different set-ups from Hurdles (combination of nucleus sampling p={0.6, 0.9}, and generation conditioning on {predicted, random} passages) and three different set-ups from WebGPT. The human-written answers are gold answers from the ELI5 subreddit for comparison with Hurdles answers, and human demonstrations for WebGPT answers.

We describe the different LFQA systems developed by prior works, which are included in comparisons used for evaluating automatic metrics in Section 4.

Krishna et al. (2021) presented a state-of-the-art LFQA system which includes a passage retriever Guu et al. (2020) and an answer generation model Roy et al. (2021).

Nakano et al. (2021) proposed to fine-tune GPT-3 Brown et al. (2020a) to interact with a search engine and compose long-form answers based on the information found. The generated answers also contain a set of reference documents found online.

A.3.2 Evaluation aspects

We describe the different evaluation aspects conducted by prior human evaluation.

Krishna et al. (2021) phrased the question as “Which generation answered the question better / was more relevant to the question?” while Nakano et al. (2021) developed detailed instructions with intermediate steps for comparing two answers, and dedicated an overall rating, phrased as “how useful the answer would be to the person asking the question, all things considered”.

Krishna et al. (2021) asked the human evaluators to choose the more coherent answer and listed repetition as a trait of incoherence.The wording was (which answer) “was more coherent / had less repetition”. In Nakano et al. (2021), the instruction for coherence evaluation focuses on whether the answer makes sense, is easy to follow and is in a logical order.

Krishna et al. (2021) instructed human evaluators to judge factual correctness of answers, with no accompanying evidence documents but permission to use search engine over Wikipedia articles. In Nakano et al. (2021), the evaluation of factuality is focused on whether the generated answer could be entailed by the evidence documents and that it doesn’t hallucinate unsupported fact. Note that “faithfulness” to the evidence articles is a different notion from the “correctness” of the answer, as the evidence articles might not always be correct or up-to-date Gao et al. (2022).

A.3.3 Example of comments mentioning different aspects for Section 3.3

A.4 Automatic Metric Implementation Details

Length statistics of the answers evaluated in 4.1 are reported in Table 13. We truncate the input if it exceeds the context window for the model. Less than 5% of the comparison data are truncated.

For each answer, we calculate ROUGE-L against the set of reference answers from ELI5 and use the maximal ROUGE-L.

We use the default roberta-large model for Englishhttps://github.com/Tiiiger/bert_score and report the maximal F1 BERT score against the set of reference answers.

We use the BLEURT-20 checkpoint as recommended and report the maximal BLEURT score against the set of reference answers.

We calculate Self-BLEU by regarding one sentence as hypothesis and all others in the same answer paragraph as reference. We report self-BLEU-5 as a measure of coherence.

We use the Stanza toolkit Qi et al. (2020) for word tokenization.

Given a question qq and an answer paragraph aa, we estimate p(q∣a)p(q|a) by computing the average log-likelihood of the question tokens conditioned on the passage using T0. Following previous work Sachan et al. (2022), we append a natural language instruction “Which question does this passage answer?” to the answer, denoted as a′a^{\prime}.

where Θ\Theta denotes the parameter of the language model and ∣q∣|\boldsymbol{q}| denotes the number of tokens in the question.

We use the BART model finetuned on the CNN/DM dataset (facebook/bart-large-cnn).

Given a question qq and an answer paragraph aa, we first encode them through the RankGen encoder, which projects them to fixed-size vectors (q,a)(\boldsymbol{q},\boldsymbol{a}). We then determine their relevance by calculating the dot product between the two vectors q⋅a\boldsymbol{q}\cdot\boldsymbol{a}. We use the T5-XXL (11B) encoder trained on both in-book negative and generative negatives.

QAFactEval Fabbri et al. (2022) is a recently proposed QA-based metric that has shown superior performane on several summarization factuality benchmark Laban et al. (2022); Maynez et al. (2020). The pipeline is carefully chosen from extensive experiments on various combinations of components in the QA-based metrics. The final pipeline consists of (1) NP from SS as Ans(S)Ans(S) (2) BART-large Lewis et al. (2020) as QGQ_{G} (3) Electra-large Clark et al. (2020) as QAQ_{A} and (4) learned metrics LERC Chen et al. (2020) as Sim(pi,si)Sim(p_{i},s_{i}). They further include an answerability classification module to determine if the question is answerable given the document DD. We report the LERC, which uses the learned metrics to compare AnsSAns_{S} and AnsDAns_{D}(a) and shows better performance compared to other metrics in our initial experiments.

A.4.1 Learned Metrics

We use pytorch-transformers Wolf et al. (2019) to implement our models. We use Quadro RTX 8000 GPUs to train our model.

We use longformer-base, consisting of 149M parameters. The training batch size is set to 16, with the initial learning rate as 1e−51e-5. We used AdamW optimizer and a linear learning rate schedule. We train the model for 5 epochs and report the result of the checkpoint with best validation accuracy. The training takes less than 5 hours with 4 GPUs.

We use the API to fine-tune the model with a batch size of 64 and a learning rate multiplier 0.05 for six epochs. Fine-tuning text-curie001 model for each epoch on OpenAI cost 11.Wedidnotusethelargertext−davinci−002model,whichwouldhavecost11. We did not use the larger text-davinci-002 model, which would have cost110 per epoch.

A.4.2 GPT-3 Two-shot

We conduct a pilot study on prompting GPT3 text-davinci-003 for the pair-wise answer evaluation task on a subset of our expert annotation data.

For each domain that has multiple experts (i.e., biology, physics, economics, and history), we evaluate on the questions for which all experts agreed on the label of the preferred answer. We randomly choose two question-answer sets as the in-context example and prompt the model on the rest of the question-answer sets. The prompt has the following format:

BETTER ANSWER: ANSWER1 (or ANSWER2) is better.

For each question-answer set, we sample three times with top pp = 1 and temperature = 0.7 to evaluate model’s consistency. The results are reported in Table 11.

Results are report in Table 11. The model is mostly self-consistent.Model also aligns with human on this small set of data where human have perfect agreement with each other, model aligns with human performance, despite variance across different domains. We leave further investigation on utilizing large language model for automatic evaluation on long-form question answering to future work.