ExpertQA: Expert-Curated Questions and Attributed Answers

Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, Dan Roth

Introduction

As the influence of large language models (LLMs) grows beyond the computer science community, experts from various fields are rapidly adapting LLMs for assistance in information-seeking scenarios. For example, medical professionals are using these systems for performing differential diagnosis Lee et al. (2023) and researchers are using them for faster literature surveys Krenn et al. (2022); Birhane et al. (2023); Owens (2023). While the use of LLMs in specialized domains has many potential benefits, it also carries significant risks. False or hallucinated claims that are confidently phrased can potentially mislead experts and propagate societal harms, especially in high stakes domains such as medicine or law Evans et al. (2021); Dash et al. (2023); Volokh (2023). Such claims can cause frustration and distrust in AI tools among individual experts in the mildest case, and propagation of misinformation and unsafe practices, in the worst case.

Providing citations or attributions within generated responses is a promising direction for alleviating such concerns. However, the quality of these attributions in model-generated responses, as well as the factuality of responses, is understudied in domain-specific settings. This is partly because we do not completely understand the specific information-seeking needs of experts. Although experts from different fields are naturally best suited to aid with such an evaluation, expert evaluations are rarely conducted, as bringing experts in the loop can be time-consuming and costly.

To bridge this gap, we construct ExpertQA, a benchmark of information-seeking questions curated by experts from 32 diverse fields. ExpertQA includes questions relevant to each field, as well as judgements from experts about how well several state-of-the-art systems perform along various axes of factuality and attribution. Having experts in the loop allows us to model a realistic information-seeking scenario that helps us understand how people in different fields use LLMs and where their capabilities fall short.

ExpertQA is constructed by asking qualified experts to formulate questions from their field that they are curious about or have encountered in their professional lives (§2.1). Responses to these questions are collected from a set of LLM-based systems that produce attributions for their answers (§3). These include purely generative, retrieval-augmented, and post-hoc attribution systems. We then ask experts to validate the claims and evidences found in the responses to their own questions (§2.2). Experts judge each claim for its informativeness to the question, its citeworthiness, and factuality. They are also asked to judge how faithful the claim is to an accompanying evidence and rate the reliability of the evidence source. Finally, experts revise each claim so it is faithful to a reliable source and make a best effort attempt at ensuring the claim is factual. This overall process is described in Figure 1.

We first use ExpertQA to evaluate representative systems from which responses are sampled (§4). Our findings suggest that:

Retrieve-and-read systems often generate complete attributions compared to LLM prompting and post-hoc attribution, but struggle to produce citations for all cite-worthy claims.

The retrieval source significantly impacts the quality of attribution and overall factuality.

High-stakes domains such as medicine and law suffer from a large percentage of incomplete attributions (35% and 31% incomplete attributions respectively) and many attributions come from unreliable sources (51% attributions are rated as not reliable by experts).

We also measure the extent to which existing automatic methods for attribution and factuality estimation Bohnet et al. (2022); Min et al. (2023) correlate with expert judgements (§5). We find that these metrics fall short in correlating with reference judgements of attribution and factuality. However, adapting these metrics to our data through finetuning results in improvements across domains.

The revised answers we collect can be used for improving and evaluating future models on long-form question answering. While similar datasets have been proposed Fan et al. (2019), examples in ExpertQA are more realistic and contain verified answers edited by experts. Furthermore, unlike existing datasets, questions in ExpertQA contain few vague questions because they include questions professionals have encountered in their practice. We establish several baselines and show that we can improve models by finetuning on ExpertQA but that there is substantial room for improvement, both in terms of ROUGE and QAFactEval (§6).

ExpertQA: Annotation Tasks

The annotation is conducted in multiple longitudinal stages and we describe each of these below. In the first stage, we ask experts to write questions from their field (§2.1). In the next stage, we present responses sampled from various systems back to the same experts for analysis (§2.2). Further details about annotator backgrounds, annotation cost and screenshots of our interface, are presented in Appendix A.

Participants are recruited through Prolific and are qualified as experts if they have attained a formal education in the field and have worked in the field for at least 3 years. They are first asked to write questions from their field of expertise. They are told that this question could be one they have encountered in their profession or one they are curious about. We ask them to formulate challenging technical questions, for which it may not be possible to find a single document on the web that answers them completely.

Each expert is asked to write 5 questions and to specify the question type(s) for each question (as shown in Table 2). We formulate these question types broadly based on existing work that attempts to classify information needs Rose and Levinson (2004). Because of their practical nature, at least two of the questions are required to be scenario-based questions (Type V, Table 2). We collect more than 3000 questions this way from 524 experts across 32 fields. We manually filter all these questions for coherence and relevance to the field and our initial question corpus contains 2507 questions. Examples of these questions from different fields are presented in Table 1.

2 Stage 2: Answer and Claim Annotation

Next, we generate responses for the questions from stage 1 by prompting six different systems that provide attributions with their answers. These systems are described in §3. We split each answer into claims, where claims are considered at the granularity of a sentence and extracted using the spaCy sentence tokenizer Honnibal and Montani (2017).We also considered further increasing the atomicity of claims (like Kamoi et al. (2023)) but finer-grained atomic claims incur considerably higher annotation cost.

In this stage of annotation, experts validate responses to their own questions. This is beneficial as experts are best qualified to evaluate answers to their own questions. We noticed a low attrition rate in our study as around 92% of annotators from stage 1 validated at least 1 of their own questions in stage 2. Since this task is intensive, a single annotation task is broken down into 1-3 question-answer pairs. The following properties of answers and claims are evaluated, and are presented to annotators in the same order as below. Properties that judge answer quality are marked with and those that judge evidence quality are marked with .

Participants are first asked to judge whether the complete answer is useful in answering the question. Usefulness is measured based on whether the answer is at least partially responding to the question. Usefulness is marked on a scale of {useful, partially useful, not useful at all}.

( + ) Claim / Evidence Attribution.

Attribution is judged based on whether a claim is supported by its accompanying evidence, following a similar design as Rashkin et al. (2021); Bohnet et al. (2022). Support may be judged as complete, partial or incomplete. If no evidence is provided, support needs to be marked Missing and if the evidence is inaccessible, it needs to be marked N/A. Annotators are told that they can assume that certain common sense facts don’t need to be explicitly stated in the evidence to judge support. If the evidence included multiple documents, annotators judge support for the claim collectively using all documents. Taking inspiration from Kamoi et al. (2023), if the claim is partially supported, annotators are asked to specify the unsupported span and the reason why it is unsupported.

Evidences can come in the form of URLs or passage evidences, depending on the system. Judging attribution can be difficult for passage evidences as relevant context might be missing, so we provide URLs along with attributed passages for context. Annotators are instructed to only use the passage for judging attribution in these cases.

() Claim Informativeness.

To differentiate between the relevance of different claims for the question, we asked annotators to label informativeness of each claim. A claim may be judged as central to answering the question (very relevant), making a relevant point that is slightly important to answer the question (a bit relevant), making a relevant point that isn’t too relevant to answering the question (not too important) or making a peripheral point that is not relevant to answering the question (uninformative).

() Claim Factuality.

Next, we ask annotators to label their best estimate of the factual correctness of each claim. They are asked to judge factuality based on their own expertise, the evidence returned by the system, and minimal browsing on the internet if needed (lasting no longer than 2-3 minutes). This judgement is also collected on a Likert scale (Definitely correct, Probably correct, Unsure, Likely incorrect and Definitely incorrect). We ask annotators to be conservative in their judgements of factual correctness, labeling Definitely correct only if every word in the claim is correct.

() Reliability of Source Evidence.

Experts may judge certain sources as more reliable than others. For example, webmd.com is a patient-facing website for medical information that may be considered reliable by non-medical experts, but medical experts may not put as much trust into the evidence from this website. Catering attributions to domain experts will require presenting attributions that they consider as credible. Therefore, we ask annotators if the accompanying evidence (if any) is found on a website they would consider reliable (on a scale of Reliable, Somewhat Reliable, Not reliable at all).

() Worthiness of Claim Citation.

Previous work has highlighted that not all claims are equally worthy of citation Bohnet et al. (2022); Alam et al. (2023). Certain claims might be considered too obvious or fundamental to a user’s domain of expertise and hence, might not need a citation (for example, the definition of a cell for a biologist). Annotators are asked to simply state if the claim is necessary to be cited (as a binary choice). We note that the cite-worthiness of a given claim might depend heavily on a user’s prior knowledge and needs, which can vary across experts from the same field.

( + ) Claim/Evidence Revision.

After labeling the above properties, annotators edit the claim and evidences to ensure that the claim is factually correct and the given references support the claim. For instance, if the claim is false or uninformative, annotators could delete it. Further, if the evidence is incorrect or insufficient, they could remove it. While doing so, we did not require them to replace missing/incorrect evidence with correct evidences because doing that can be too time-consuming.

Systems Evaluated

We now describe the classes of systems from which we sampled responses for evaluation (also see Figure 3). All systems produce an answer string and attributions in the form of in-line citations for the sentences in the answer. Attributions are either returned as URLs to webpages or passages along with URLs to webpages from where they are retrieved. Further experimental details and prompts are included in Appendix B.

In this paradigm, we prompt large language models in a closed-book fashion Brown et al. (2020); OpenAI (2023) to generate an answer with in-line citations where they provide URLs for each citation. Note that this is unlikely to work as the model essentially has to generate a URL from its parametric memory. Nevertheless, we consider GPT-4 as the LLM from which we sample responses (gpt4). The prompt used is given in Table 9.

Post-hoc retrieval.

This system differs from the above, as we only prompt LLMs to generate answers without attribution, and perform retrieval of evidence for a claim as a post-hoc step. This renders the attributions naturally unfaithful, but we believe this is still a worthwhile approach to investigate because of the strength of LLMs as generators and retrievers independently. The attribution corpora we consider are Sphere Piktus et al. (2021) (post_hoc_sphere_gpt4), which is a large static dump of CommonCrawl, and Google search results (post_hoc_gs_gpt4).

Retrieve-and-read.

In this class of systems, we first retrieve evidence for a question and then generate an answer by prompting a model to use the retrieved evidence to answer the question. As our attribution corpus, we again consider Sphere Piktus et al. (2021) (rr_sphere_gpt4) as well as Google search results (rr_gs_gpt4). We use sparse retrieval using BM25 Robertson et al. (2009) for retrieving from Sphere. We then generate an answer using GPT-4, by including the retrieved evidence as context in our prompt. The answer generator is instructed to also generate in-line citations for each sentence, which refer to the passages provided in the context. The prompt used is given in Table 10.

Commercial systems.

We also consider commercial systems that have recently gained popularity, such as BingChat. The precise implementation of these systems is proprietary, but we can still draw conclusions about the utility of these systems for domain experts through our human evaluation. In this work, we sample responses using the balanced mode of BingChat (bing_chat).

1 Response Sampling

We sample uniformly from all systems but exclude all abstained answers and constrain the number of claims for each answer to be at most 10 for lowering annotation costs. During annotation, we noticed that attributions from gpt4 often pointed to broken links and evaluating such attributions is not meaningful, so we sampled fewer responses from that system. The number of examples included from each system and the abstention rate of the systems (how frequently a system responds with one of the predefined strings indicating it cannot provide an answer) is presented in Table 3.

Analysis

The total number of examples validated in stage 2 of our annotation is 2177. The average number of claims and average number of tokens across these examples is 5.79 and 152.12 respectively (the overall distributions are presented in Figure 4). The distribution of examples across fields is presented in Figure 2. A large percentage of examples in ExpertQA come from high-stakes fields such as Healthcare/Medicine and Law. Table 2 presents the distribution of questions across different question types shown to annotators in stage 1. The largest number of questions are Type V because participants were asked to write 2/5 questions of this type. After those, the largest number of questions are open-ended questions (Type I) and directed questions (Type II).

2 Manual Analysis

To estimate the agreement of human labels in ExpertQA, we (the authors) labeled our agreement with the reference labels from two fields to which the authors collectively belong. Specifically, we sampled 60 questions from Engineering and Technology and another 60 from Healthcare / Medicine, where we included 10 questions from each system. For each claim in the responses to these 60 questions, we label binary agreement with the annotated properties from §2.2.

Our analysis showed fairly high agreement (> 85%) for most labels in both fields. The agreement for attribution labels for medical claims was found to be slightly lower than engineering claims. We noticed that medical claims allow for more nuance in interpretation, compared to more objective claims in engineering.

3 Results

We present the Likert distribution for claims across all systems and properties in Figure 6. Below we summarize our main conclusions from the analysis.

We find that ∼\sim87-89% of answers from gpt4 are marked useful. The retrieve-and-read systems (as well as bing_chat, which also retrieves evidences first) are marked slightly less useful (73-80%), likely because retrieved evidences are not always highly relevant. Choosing relevant evidences using Google search results in more useful answers than with the smaller Sphere corpus and a sparse retriever.

Retrieve-and-read systems have a stronger inductive bias to use the retrieved evidence to answer the question. However, they do not always produce attributions for cite-worthy claims (18% of these claims are missing attributions)Figure 6 shows the Likert distribution of attribution labels on those claims deemed cite-worthy by experts.. On the other hand, post-hoc attribution systems return attributions for every single claim, by definition, but return more incomplete attributions. Lack of context during post-hoc retrieval can be an issue for retrieving valid attributions. For example, a claim such as "Targeted therapies, such as PARP inhibitors or CDK4/6 inhibitors, may be useful depending on the specific genetic makeup or molecular features of the tumor ." does not actually contain any details about the fact that the question asked about breast cancer, which can make it hard to find relevant retrievals. Therefore, we might need approaches that can make claims standalone, such as Choi et al. (2021), to improve post-hoc citation quality.

Finally, without retrieval, we found that gpt4 often generates citations to URLs that link to plausible & trustworthy domains (for eg, nasa.gov for astronomy and nih.gov for medical claims), but the content on these webpages is often totally mismatched (more than 60% of the time).

Both vanilla prompting and retrieval-augmented systems generate mostly very relevant claims to the question.

At the same time, a significant percentage of claims (30-40) are not very relevant. This may include void claims (that simply restate the question or state simplistic facts). This suggests that there is a lot of room in making answers concise and relevant. In addition, using Sphere as the retrieval corpus results in fewer informative claims than Google search for retrieve-and-read systems.

Just over half the claims are labeled as definitely correct by experts.

While a significant percentage of claims are labeled as correct (probably or definitely), experts do not instill high confidence in the factual correctness of claims. This might be because it is hard to judge factuality with a high degree of confidence, in a short time frame. Once again, a smaller retrieval corpus and retriever (rr_sphere_gpt4) results in less factual claims as the model may be more likely to hallucinate.

The retrieval corpus has a significant effect on expert judgements of source reliability.

Expert judgements of reliability are directly influenced by the credibility of the sources from which evidences are retrieved. Corpora such as Sphere, which do not account for source, often present evidences that are unreliable to experts (for both rr_sphere_gpt4 and post_hoc_sphere_gpt4). For example, in a question about breast cancer, evidence from a comment on a blog is retrieved and is naturally considered Not reliable at all by the expert. Using Google search as the retrieval system, improves source reliability judged by experts.

Majority of claims are deemed cite-worthy across systems.

Only around 17-22% claims are judged not citeworthy by the experts. This suggests that most claims in responses to expert-curated warrant providing supporting evidence. Non-cite-worthy claims are mostly either uninformative or too basic for an expert from the field.

Effect of Retrieval Corpus.

While sampling responses, we instruct models to abstain if they cannot answer the question faithfully and accurately. Among retrieve-and-read systems, we find that the abstention rate is significantly higher for systems that use Sphere as the retrieval corpus compared to Google search. Across claim properties, we find that the retrieval corpus has a significant impact on human judgements. Systems that use Sphere as the retrieval corpus instead of Google search result in more missing attributions, fewer correct and informative claims, attributions with more evidences from unreliable sources, and overall less useful answers.

Domain and Question Type Trends.

Figure 10 shows the distribution of labels across all fields. There are few clearly discernible patterns from the trends across fields. The percentage of claims labeled as being Definitely or Probably correct is fairly high (>85%) for many fields. However, we note that across all annotated claims, high-stakes domains such as medicine and law can suffer from a significant percentage of incomplete attributions (around 35% of medical claims and 31% of legal claims worthy of citation are not supported with appropriate evidence). Further, a large percentage of claims present evidences from unreliable sources for these domains (for eg, ∼\sim51% of medical claims have attributions from sources that are not labeled Reliable by experts).

Across question types, systems clearly struggle with Type VI questions that request for a list of resources, as claims are less informative, factual, and supported by evidence. Other question types that appear to be challenging are Type IV and Type VII, which seek advice for a problem or opinions on a topic respectively.

Automatic Estimation of Attribution and Factuality

Next, we study the effectiveness of current automatic systems for attribution Honovich et al. (2022) and factuality (correctness) estimation Min et al. (2023) in the context of ExpertQA. In both cases, we observe that current systems show high precision but low recall when compared with human judgements of attribution and factuality.

Under the framework of attributable to identified sources (AIS) Rashkin et al. (2021), i.e. judging whether model generated content can be verified against given attributions, previous work has found NLI models to be effective in providing automated AIS (AutoAIS) estimations (Gao et al., 2023; Bohnet et al., 2022).

To understand the effectiveness of AutoAIS on ExpertQA, we follow the settings in previous work and use the NLI classifierhttps://huggingface.co/google/t5_xxl_true_nli_mixture from Honovich et al. (2022) as an AutoAIS system to predict the attribution labels of claim-evidence pairs in ExpertQA. The model gives a binary decision for whether a hypothesis claim is entailed by the evidence as premise. For evidence longer than the maximum sequence length of the model (512), we apply the retrieve-and-rerank technique from Schuster et al. (2022), where we split evidence into sentences and take top-2 sentences with highest entailment confidence predicted by the NLI model as evidence.

Table 4 shows the macro-averaged AutoAIS scores for the set of claims marked as having attribution with complete support from each system. Compared to our human judgments, the AutoAIS estimations show a large variance in terms of the averaged completeness of attribution across systems. Notably, attributions generated by post-hoc retrieval receive much lower AutoAIS scores, whereas retrieve-and-read systems get higher scores.

We compare the per-claim AutoAIS predictions to human judgement of attribution completeness in Table 5. The results suggest that AutoAIS estimations with a NLI model overall have high-precision yet low-recall against human judgements of attribution completeness. To understand the discrepancy between NLI model behavior vs. human judgement of attribution completeness, we highlight a few typical examples of attribution(s) in Table 6. For NLI, every part of the hypothesis or claim is expected to be directly verifiable against the premise, while we observe during attribution verification, human judgements involve more implicit world knowledge, e.g. calcium carbonate is an alkali. Another common type of mistake from AutoAIS involves combining information from multiple pieces of evidence. We observe mutli-source attributions to be particularly common among bing_chat and retrieve-and-read systems.

2 Automatic Factuality Estimation

Prior work has proposed methods Manakul et al. (2023); Min et al. (2023) to automatically estimate the factuality of model-generated responses. In particular, we use and study FActScore Min et al. (2023) as an automatic estimator for the factuality of claims. We first break down each claim into finer-grained atomic claims using few-shot prompting with text-davinci-003, using the same prompt and setting as Min et al. (2023). We retrieve the top-k (k=3) relevant passages using Google search using an atomic claim as the query. The evidence paragraphs for each atomic claim are then included in a prompt sent to gpt-3.5-turbo, along with the atomic claim, where the model is instructed to say whether the atomic claim is True or False. The final factuality score for a claim is then calculated by averaging the scores of all the atomic claims.

In Table 7, we report the F1 scores of the factual (T) and non-factual (F) classes and the micro-averaged overall F1 scores of the FActScore factuality scores and the reference factuality labels. FActscore scores are thresholded at 0.5 to get binary scores and reference factuality labels are 1 if the claim’s factuality is labeled as Probably correct or Definitely correct, and 0 otherwise.

We find that automatic factuality estimation struggles to identify non-factual claims in our dataset. In particular, predicted labels have low recall of non-factual claims. This is more often the case for retrieve-and-read systems, where the answer is generated based on retrieved evidences. The other systems use GPT4’s parametric knowledge for answer generation, which might make it slightly easier for an evaluator like ChatGPT to judge factuality.

Long-form QA Evaluation

A beneficial output of our annotation pipeline is the revised answers produced by annotators. These answers are vetted by experts to contain factual information and compose a new long-form QA dataset, ExpertQA. We consider two types of splits for ExpertQA (both 80-10-10): a random split of the data and a domain-wise split, where 80% of a field’s data is included in the training set and 10% is included in both validation and test sets.

We consider standard metrics used for evaluating long-form question answering, that are based on similarity to a reference answer, i.e., ROUGE Lin (2004) and those focused on evaluating factual consistency through question-answer pairs generated on a reference answer, i.e., QAFactEval Fabbri et al. (2022).

2 Baselines

We finetune the following open-source language models: FlanT5-11B Chung et al. (2022), Alpaca-7B Taori et al. (2023), Vicuna-7B Chiang et al. (2023) and LLaMa2-7B-Chat Touvron et al. (2023). We finetune these models with the same prompts as the ones used in their training (provided in Tables 11, 12). Further, we also report results with Llama2-70B-Chat without finetuning (marked *).

3 Results

Our results are shown in Table 8. We find that both Llama2-7B and Vicuna-7B outperform FlanT5-11B despite the smaller model size, likely due to additional instruction finetuning for both those models. We observe that finetuning significantly improves performance (results without finetuning are in Table 13), and Llama2-70B performs worse than finetuned systems under zero-shot prompting.

Related Work

With the rapid adaptation of large language models, a few different classes of systems have been proposed for generating attributions for the responses they produce. The first class of systems is vanilla LLM prompting Tay et al. (2022); Weller et al. (2023), where LLMs such as GPT-4 OpenAI (2023) are prompted to return attributions (in the form of titles of references, optionally accompanied with URLs) along with their answers. Since these systems are not explicitly trained to produce attributions, they can often hallucinate the references they provide Agrawal et al. (2023). Another class of systems is retrieve-and-read systems Guu et al. (2020); Borgeaud et al. (2022); Izacard et al. (2022), which first retrieve evidence relevant for a query, and then generate an answer based on the retrieved evidence. These systems are sometimes trained on demonstrations of humans browsing for information Nakano et al. (2021); Thoppilan et al. (2022); Menick et al. (2022), which allows them to jointly generate answers and citations. Finally, post-hoc retrieval Gao et al. (2023); He et al. (2022) involves retrieving attributions after answering a query using both the query and response for retrieval, and optionally revising the answer based on the attribution. For sampling responses, we consider all three classes of systems described above.

Attribution Analysis

Prior work has conducted analysis of attributions produced by systems in response to queries from existing NLP datasets Rashkin et al. (2021); Bohnet et al. (2022); Dziri et al. (2022); Liu et al. (2023); Muller et al. (2023). Notably, Rashkin et al. (2021) propose the framework of Attributable to Identified Sources (AIS) for performing human evaluation of attributions. Analysis from these works suggests that systems are still far from providing precise attributions with sufficient recall for all citeworthy statements. In our work, we recognize that this is problematic in specific domains where attribution precision and recall are both critical.

Automatic methods to measure attribution have also been explored. Briefly, attribution has been automatically measured through the use of textual entailment models Bohnet et al. (2022); Yue et al. (2023), by checking the consistency of answers produced to the same question using the evidence or the response Wang et al. (2020) and prompting LLMs and finetuning LLMs on tasks such as NLI relevant for judging attributions Yue et al. (2023).

Previous efforts at collecting gold attribution data have been conducted by repurposing Wikipedia citations Kamoi et al. (2023); Petroni et al. (2022), or existing datasets like Natural Questions Bohnet et al. (2022); Kwiatkowski et al. (2019), and relying on large-scale human annotation Chen et al. (2022); Dziri et al. (2022); Kamalloo et al. (2023). We note that with regards to the information needs of experts, this data can be restricting in terms of the accompanying corpus, and contain limited complexity when it comes to expert-curated content (for instance, experts are likely to use and trust sources other than Wikipedia).

Factuality Analysis.

Analysis of factuality or truthfulness of LM generations has been conducted extensively in prior work Thorne et al. (2018); Evans et al. (2021); Maynez et al. (2020); Pagnoni et al. (2021); Lin et al. (2021); Muhlgay et al. (2023). Factuality is closely linked to work studying hallucinations Ji et al. (2023), which includes closed-domain hallucination Kryscinski et al. (2020); Maynez et al. (2020), where hallucinations are analyzed in reference to some context, as well as open-domain hallucination (for example, Manakul et al. (2023)). The factuality labels we collect as part of ExpertQA elicit a best-effort judgement of truthfulness of claims from experts.

Prior methods proposed to verify the factuality of statements include zero-resource approaches Manakul et al. (2023); Kadavath et al. (2022); Agrawal et al. (2023); Azaria and Mitchell (2023); Min et al. (2023), which operate mostly in a black-box fashion, and resource-enriched approaches Thorne et al. (2018); Guo et al. (2022); Chen et al. (2023); Feng et al. (2023), that verify statements by checking their validity against external databases or resources. We use one such recent method called FactScore Min et al. (2023) to evaluate how well the factuality labels in our dataset correlate with automatic factuality judgements. Min et al. (2023) split claims into finer-grained atomic claims and using LMs to verify these atomic claims. Dedicated benchmarks to study hallucination have also been proposed Liu et al. (2022); Li et al. (2023), but these are often synthetically generated and not based on real user queries and responses to them.

Long-form QA.

Existing long-form QA datasets are often created using naturally occurring sources on the web such as search queries Nguyen et al. (2016); Stelmakh et al. (2022) and forums Fan et al. (2019). Several issues have been pointed out with conducting robust evaluation for long-form QA Krishna et al. (2021). Previous work has also suggested that evaluating finer-grained claims provides higher agreement of factuality among annotators Krishna et al. (2023). Keeping this in mind, we construct ExpertQA to cover practical information needs of experts and collect fine-grained judgements of factuality along with vetted answers.

Xu et al. (2023) also conduct human evaluation of long-form answers with experts from 7 fields. Notably, they emphasize the importance of evaluating multiple aspects of long-form answers, which are also considered in our work.

Domain-specific QA.

Several domain-specific datasets for question answering have also been proposed in prior work. Specifically, existing work has presented datasets for the medical domain Tsatsaronis et al. (2015); Pampari et al. (2018); Jin et al. (2019, 2021); Pal et al. (2022), legal Guha et al. (2023), technology Dos Santos et al. (2015) and others. There is also work that proposes datasets with examples spanning multiple domains Rogers et al. (2020); Reddy et al. (2019); Hendrycks et al. (2021). However, these datasets often have a very limited coverage of domains. Importantly, most of these are not created with experts-in-the-loop and hence may not always represent natural information seeking distributions, that go beyond factoid questions. Further, they do not contain long-form answers and attributions for the answers. Finally, because of large-scale pretraining, we are aware that many such datasets are under the risk of contamination Sainz et al. (2023). Hence, we choose to crowdsource questions from experts, and further also get annotations for attribution and factuality.

Conclusion and Future Work

Our study suggests that although large language models show a lot of promise for aiding domain experts, there is large ground to cover in being able to provide expert-level guidance that is factual and also supported by reliable sources Metzler et al. (2021). While LLMs make information access and search substantially easier for domain experts, improving factuality and attribution of answers is necessary to improve trustworthiness in these systems. Repurposing general-purpose language models for domain-specific purposes requires having experts-in-the-loop, so we can understand their information needs and how models fall short in meeting them. We hope that our benchmark, ExpertQA, can benefit the community with improved methods for attribution & factuality estimation, and long-form question answering evaluation.

Limitations

In most cases, claims in our dataset are sentences that may not represent singular information units. This lack of atomicity in claims means that properties such as factuality and attribution need to be judged exhaustively for a claim. Collecting human judgements for finer-grained atomic claims can be significantly more expensive and is not explored in this work.

Claim Extraction.

Extracting sentence-level claims from a generated answer for the purpose of evaluation is performed by using a sentence tokenizer. However, we note that existing tokenizers suffer from sentence tokenization errors (for example, when lists or tables are present in answers). This resulted in a small number of claims being excessively long and hard to evaluate.

Field Coverage.

Even though we tried to cover a wide range of fields in our dataset, we missed covering questions from certain fields. Finding experts from rarer fields can be especially hard. We will consider further expanding ExpertQA to more domains, so that it can be more broadly useful. In addition, the examples in our dataset represent the information needs of English-speaking annotators primarily based in Europe, the Americas and Africa.

Subjectivity of labels.

Some of the properties of claims can elicit more subjective judgements, which can vary between experts from the same field. This subjectivity is not inherently captured in our data through multiple judgements, but we do estimate agreement using claims from engineering and medicine through our own labels (§4.2).

Acknowledgements

First, we would like to thank the 484 annotators who took the time and effort out to help out with this study. We would also like to thank Artemis Panagopoulou, Alyssa Hwang, Nelson Liu and Chris Alberti for helpful comments and discussions.

References

Appendix A Annotation Details

The 484 participants involved in our study came from 26 different countries, across Europe, Africa, Oceania, North and South America. The participants were recruited through Prolific, a crowdsourcing platformhttps://www.prolific.co. To qualify as experts, participants were required to have attained a formal education in the field and have worked in the area for at least 3 years. Participants were told that their annotations will be used to evaluate the capabilities of large language models to provide truthful answers with well-supported evidences to questions from different fields. They were also informed that the data will be released publicly upon the completion of the study.

Annotator fields.

The initial set of fields were listed by going through university department names, and ensuring that we cover a wide range of disciplines. Upon completing stage 1 of our annotation, we further refined these fields to represent a diverse set, for which we have enough experts.

Annotation costs.

In both stage 1 and stage 2, annotators were compensated at the rate of $15 per hour with additional bonuses when annotators spent more time than we anticipated. The average time taken for stage 2 annotations was 13.83 minutes per question-answer pair, and there was a lot of variance in time taken across annotators.

Annotation Interface.

Figure 7 and Figure 8 show screenshots of our stage 2 annotation interface in the order the task was presented to annotators.

Appendix B Experimental Details

Across all systems, for generating responses from gpt4, we use a temperature of 1.0, and a maximum length of 2048 tokens. For all retrieval componenets, we use text-embedding-ada-002 as the embedding model. The retrieve-and-read systems first retrieve top-k (k=5) evidence passages from Sphere or top-10 Google search results using the question as the retrieval query. Google search results are split into passages of 1000 tokens with 200 tokens of overlap between subsequent chunks.

On the other hand, the post-hoc citation systems simply use the claims from gpt4 responses, but generate their own attributions by retrieving evidence for each claim in the answer. Post-hoc retrieval systems use the top-k passages (k=5) retrieved from Sphere or the top-10 Google search results with the claim as the retrieval query. Search result are split into passages the same way as retrieve-and-read systems.

Automatic attribution and factuality estimation.

For automatic attribution with AutoAIS, we use the t5_xxl_true_nli_mixturehttps://huggingface.co/google/t5_xxl_true_nli_mixture with 11B parameters by Honovich et al. (2022). For finetuning the t5_xxl_true_nli_mixture model on the train split of ExpertQA, we use the DeepSpeed ZeRO optimization Rajbhandari et al. (2020) with stage 3, a batch size of 1, a learning rate of 1e−41e^{-4} and train models for 3 epochs.

Long-form QA.

For finetuning FlanT5-11B, we use a batch size of 2, maximum sequence length of 512, a learning rate of 1e-4 and train models for 3 epochs. For finetuning both Llama2-7B and Vicuna-7B, we use a batch size of 4, maximum sequence length of 2048, learning rate of 2e−42e^{-4} and train models for 3 epochs.

B.2 Prompts

The prompts used to generate responses from gpt4 and bing_chat is provided in Table 9, while the prompt used to generate responses for retrieve-and-read systems is in Table 10.

For factuality estimation, we use the same prompts as Min et al. (2023) for both claim decomposition and atomic claim factuality prediction. Finally, for long-form QA baselines, we use the prompt in Table 11 for Llama and Table 12 for Vicuna.

Appendix C Additional Plots

Examples from all fields included in ExpertQA are shown in Table 14. We show the distribution of all question types (from Table 2) across all fields that are part of ExpertQA in Figure 9.

In Table 10, we summarize the label distribution of all claim properties across fields and in Table 11, we summarize the label distribution of all claim properties across question types.

In Table 13, we summarize results on long-form QA before and after finetuning models on both ExpertQA splits.