FreshLLMs: Refreshing Large Language Models with Search Engine Augmentation
Tu Vu, Mohit Iyyer, Xuezhi Wang, Noah Constant, Jerry Wei, Jason Wei, Chris Tar, Yun-Hsuan Sung, Denny Zhou, Quoc Le, Thang Luong
Introduction
Recent large language models (LLMs) such as Bard and ChatGPT/GPT-4 https://bard.google.com, https://chat.openai.com are designed to be versatile open-domain chatbots that can engage in multi-turn conversations on diverse subjects. Despite their impressive capabilities, these LLMs often “hallucinate” plausible but factually incorrect information (Maynez et al., 2020; Liu et al., 2023b), which reduces the trustworthiness of their responses, especially in settings where accurate and up-to-date information is critical. This behavior can be partially attributed to the presence of outdated knowledge encoded in their parameters. While additional training using human feedback (Ouyang et al., 2022) or knowledge-enhanced tasks can mitigate this issue, it is not easily scalable for real-time knowledge updates (e.g., stock price of a company). In-context learning (Brown et al., 2020) is an appealing alternative in which real-time knowledge can be injected into an LLM’s prompt for conditioning generation. While recent work has begun to explore augmenting LLMs with web search results (Lazaridou et al., 2022; Press et al., 2022), it is unclear how to take full advantage of search engine outputs to increase LLM factuality.
In this work, we collect a novel QA benchmark, dubbed FreshQA, to evaluate the factuality of existing LLMs. FreshQA consists of 600 natural questions that are broadly divided into the four main categories shown in Figure 1. FreshQA’s questions span a diverse set of topics with diverse difficulty levels (requiring single-hop and multi-hop reasoning), and require a model to “understand” the world’s up-to-date knowledge to be able to answer correctly. Additionally, FreshQA is dynamic in nature: some of the ground-truth answers may change over time, and a question classified under a specific category may undergo reclassification at some later point in time (e.g., the current false-premise question “How long has Elon Musk been married to his current spouse?” will fall into the fast-changing category if Elon Musk gets married again in the future).
We benchmark how well different LLMs perform on FreshQA by prompting them with questions and optionally a few question-answer demonstrations and then sampling a response. Then, we conduct an extensive human evaluation of the factual accuracy of the models’ responses, consisting of more than 50K judgements. We evaluate each response in a two-mode evaluation procedure: Relaxed, which measures only whether the main answer is correct; and Strict, which measures whether all of the claims in the response are factual and up-to-date (i.e., no hallucination). Our study sheds light on the factuality of old and new LLMs and reveals different model behaviors across question types. Unsurprisingly, there are flat scaling curves on questions that involve fast-changing knowledge: simply increasing the model size does not lead to reliable performance gains. We also observe similar trends on false-premise questions, though several LLMs are able to debunk a false-premise question if explicitly asked “Please check if the question contains a valid premise before answering”. Overall, FreshQA is challenging for current LLMs and leaves ample room for improvement.
Motivated by these findings, we further investigate how to effectively improve LLMs’ factuality by grounding their responses to accurate and up-to-date information from search engines. Given the rapid development of ever larger LLMs and the ever-changing nature of knowledge, we explore in-context learning approaches that allow an LLM to attend over knowledge provided at inference time through its prompt. We develop FreshPrompt, a simple yet effective method that, for a given question, takes full advantage of a search engine by extracting all up-to-date and relevant information (including knowledge from relevant questions that search users also ask) and uses few-shot in-context learning to teach a model to reason over retrieved evidences and figure out the right answer. We show that FreshPrompt significantly boosts LLMs’s factuality: for example, our best GPT-4 + FreshPrompt variant yields an improvement of 32.6% and 49.0% accuracy over the vanilla GPT-4 on FreshQA under Relaxed and Strict, respectively. Since our method requires no additional training, it is flexible and applicable to a variety of scenarios.
Taken together, our key contributions include:
We introduce a novel dynamic QA benchmark, FreshQA, which features a diverse set of question and answer types, including questions whose answers may change over time and questions whose premises are factually incorrect. We make our dataset freely available and commit to updating the ground-truth answers at a regular schedule to encourage exploration of methods to improve LLMs’ factuality.
We benchmark a wide range of both closed and open-source LLMs on our dataset. Through an extensive and rigorous human evaluation study, we shed light on limitations of current LLMs: they struggle on fast-changing, false-premise, and multi-hop questions, and our two-mode evaluation captures increased hallucinations produced by techniques such as chain-of-thought prompting (Wei et al., 2022).
We present FreshPrompt, a simple in-context learning method that can substantially boost an LLM’s factuality compared to competing search-augmented approaches by effectively incorporating factual and up-to-date information from a search engine into the model’s prompt. Furthermore, we perform a series of sensitivity and ablation analyses to better understand what facets of FreshPrompt contribute to its success.
FreshQA
In this section, we address the growing need to assess LLM factuality by curating a novel QA benchmark, FreshQA, with 600 questions that cover a wide spectrum of question and answer types.
We collected FreshQA by recruiting both NLP researchers (including the authors and their colleagues) and online freelancersWe use Upwork (https://www.upwork.com) with a compensation rate of $2 per example. to write questions of varying difficulty levels and topics whose answers may change based on new developments in the world. The annotators were shown a few exemplars of the four broad types of questions defined in Figure 1. Within each of these four categories, we ask annotators to write questions at two different difficulty levels: one-hop, where the question explicitly mentions all of the relevant information needed to answer it, and thus no additional reasoning is required (e.g., “Who is the CEO of Twitter”); and multi-hop, where the question requires one or more additional steps of reasoning in order to gather all of the relevant information needed to answer it (e.g., “What is the total height of the tallest building in the world?”). Annotators were encouraged to write questions that involve fresh knowledge (knowledge that has changed recently or new events) and appear natural (i.e., plausible for a real person to type into a search engine). For false-premise questions, we requested a brief explanation elucidating why the question is flawed.Additionally, the annotators were asked to include the year the answer to the question last changed and an URL to a reputable website that supports the answer.
Upon obtaining the initial dataset, we conducted multiple thorough data cleaning and quality assessments. This involved manual review of each example to ensure well-formed questions, removal of duplicates and invalid questions (e.g., too easy or controversial), and verification of answers and supporting evidence URLs. We also manually collected supplementary valid answers for each question (e.g., different names of the same person, different date formats, etc.). To facilitate future answer updates, we excluded questions whose answers are likely to change more frequently than once per week, and additionally incorporated the expected next review date for each question.
Data size and split:
The resulting dataset is divided into a test set consisting of 125 questions for each of the four broad question types (500 total examples) and a development set comprising 25 questions for each question type (100 total examples), sampled randomly within types. Additionally, 15 examples spanning different question types were extracted for demonstration purposes (i.e., for use in few-shot in-context learning), and the remaining data was discarded. The development set is reserved for future studies and not used in this paper.Although our test set is currently balanced across question types, the distribution may change over time due to reclassification of questions from one category to another.
FreshQA requires regular updates:
Our dataset has time sensitivity since the ground-truth answers may change with new developments in the world. As such, we commit to updating the dataset regularly and encourage researchers to evaluate on the latest version of the dataset, as close to the release date of the updated dataset as possible.
2 Evaluation
All model responses were evaluated by the authors in a two-mode evaluation procedure: Relaxed, which focuses solely on evaluating the correctness of the primary answer; and Strict, which additionally examines whether all of the facts in the answer are accurate (i.e., no hallucination). Overall, our setup provides both ends of the spectrum for evaluating factuality (the difference between a model’s strict and relaxed performance provides a way to measure hallucination), offering a more comprehensive and nuanced understanding of their performance.
In both evaluation modes, we credit a model’s response only if it provides a confident and definitive answer, or the correct answer can be obviously inferred from the response. The primary or final answer when standing alone must be accurate. Any additional information that is provided must not contradict the primary answer or reshape one’s perception of it. For false-premise questions, the model must point out the presence of a false premise to receive credit. For answers that involve names of entities (e.g., people), complete names or commonly recognized names are expected. Regarding numerical answers, approximate numbers are generally not accepted unless explicitly included in the ground-truth answers. Under Relaxed, we accept ill-formed responses (including those in a non-English language), as well as hallucinated or outdated information that does not significantly impact the primary answer. Under Strict, however, a response that contains any hallucination, no matter how minor, will not receive credit. Furthermore, we accept a response in Strict when the model indicates that the information might be outdated (e.g., “As of my knowledge cutoff date in September 2021”) only if it is evident that the knowledge has not changed.Note that even without access to real-time data, a model may still provide accurate answers to certain questions involving current information, potentially through random guesses or by leveraging past valid responses (e.g., for the question “Which drama series won the most recent Primetime Emmy Award for Outstanding Drama Series?”, while “Succession” won the award most recently (as of this writing), it was also the winner in 2020, so a model trained in 2021 could potentially provide the correct answer). Figure 4 in Appendix A shows specific examples of each evaluation criteria.
Inter-rater agreement and automatic evaluation:
Two authors independently evaluated a subset of 100 answers in both modes and had an agreement of 99% for Relaxed and 96% for Strict, showing that the protocol is reliable for comparing different LLMs. Additionally, to facilitate future evaluations, we develop FreshEval, a simple automatic metric that uses few-shot in-context learning to teach an LLM to judge model responses, achieving an average agreement of 96.5% with human evaluations for Relaxed and 96% for Strict. See Appendix B for details.
Pre-trained LLMs struggle on FreshQA
We use FreshQA to benchmark LLMs that do not have access to real-time data or the ability to browse the Internet for current information.With the exception of ChatGPT and GPT-4, which have access to the current date. Note that the latest versions of these models can now browse the Internet. While all LLMs (regardless of size) predictably struggle on questions requiring up-to-date knowledge, they also underperform on false premise questions. In our experiments, we simply feed individual questions as prompts into each model and decode the model’s predictions using a temperature of 0 without fine-tuning (see Appendix C for more details).
We experiment with a series of models varying in size from 770M to 540B parameters, including basic pre-trained models such as T5 (Raffel et al., 2020; Lester et al., 2021), PaLM and PaLMChilla (Chowdhery et al., 2022), optionally using Few-shot prompting (Brown et al., 2020) and Chain-of-Thought (CoT, Wei et al., 2022);As we are interested in exploring how these methods perform without being specifically designed for FreshQA, we use the 5-shot demonstrations for TriviaQA (Joshi et al., 2017) used in Sun et al. (2023). instruction-tuned models including FLAN-T5 and FLAN-PaLM (Chung et al., 2022; Longpre et al., 2023), and OpenAI’s GPT-3.5 (Ouyang et al., 2022), Codex (Chen et al., 2021a), ChatGPT, and GPT-4 (OpenAI, 2023).
1 Results and Discussion
We visualize the accuracy of different LLMs on FreshQA in both evaluation modes in Figure 2.Table 3 and Table 4 in Appendix D contain concrete numbers under Strict and Relaxed, respectively. A first obvious takeaway is that all models struggle on FreshQA: overall accuracy ranges from 0.8% to 32.0% under Strict, and 0.8% to 46.4% under Relaxed. Switching from Relaxed to Strict results in a marked decrease in accuracy for ChatGPT and GPT-4. This is mainly due to the lack of access to up-to-date information, as they produce “outdated” answers (which often start with the prefix ‘‘As of my knowledge cutoff date in September 2021”), and in many cases, “refuse” to provide an answer (e.g., “As an AI language model, I cannot provide real-time information.”). Similarly, the accuracy of PaLM (across model sizes) drops significantly under Strict. Much of this drop is due to artifacts such as conversation-like responses with unexpected special tokens (e.g., the end-of-turn [eot]), and hallucination. In contrast, Flan-PaLM and Codex exhibit minimal hallucination due to their concise and direct answers.
LLMs struggle with questions about current information:
The lack of up-to-date parametric knowledge results in dramatically degraded accuracies across models on questions involving fast-changing or recent knowledge. GPT-4 generally obtains the highest accuracy on these questions, with the exception of questions about recent knowledge (i.e., since 2022) under Strict where it underperforms Flan-PaLM and Codex, but it never exceeds 15% across both evaluation modes. Our evaluation confirms that ChatGPT and GPT-4 have been exposed to data containing information beyond their knowledge cutoff date (Appendix E). Additionally, GPT-4 is more reluctant to answer fast-changing questions (refusing to answer 60% of the time) compared to ChatGPT (16%).
Questions with false premises pose a hurdle for LLMs:
All models struggle on questions with false premises, and using larger models does not increase accuracy for T5 and PaLM (“flat scaling”), with performance within the range of 0.0% to 1.6%. GPT-3.5, ChatGPT, and GPT-4 demonstrate much superior accuracies to all other models, achieving accuracies between 25.8% to 42.7% under Strict and 32.3% to 66.9% under Relaxed. ChatGPT performs the best under Strict (42.7%) while GPT-4 is the most accurate model under Relaxed (66.9%), with an impressive accuracy of 83.9% on questions about knowledge before 2022. These results suggest that OpenAI’s models are likely trained to cope with false-premise questions.
CoT increases hallucination:
Overall, Few-shot and CoT prompting are beneficial for large models and sometimes advantageous for moderately-sized models on questions with valid premises, especially on questions about never-changing or old knowledge. Under Strict, Few-shot and CoT yields +36.1% and +26.9% respective accuracy improvement over zero-shot prompting with PaLM 540B on questions involving knowledge before 2022 (+21.9% and +29.7% under Relaxed). CoT largely demonstrates superior performance compared to Few-shot under Relaxed, whereas Few-shot obtains better results under Strict, as CoT introduces more room for hallucination.
Multi-hop reasoning is challenging for several models:
T5 Large and XL are incapable of dealing with multi-hop questions, while Flan-PaLM 540B, Codex, and GPT-3.5 suffer the most when switching from one-hop to multi-hop questions. GPT-4 remains stable across these two types of questions (with a difference of less than 2% in accuracy across settings). See Appendix D for details.
Prompting Search Engine-Augmented Language Models
The low accuracies reported in the previous section are largely unsurprising, as none of the models we evaluated had access to real-time information. In this section, we evaluate the impact of search engine augmentation to LLMs on FreshQA. We present FreshPrompt, a simple few-shot prompting method that substantially boosts FreshQA performance of an LLM by incorporating relevant and up-to-date information retrieved from a search engine (Google Search) into the prompt.
Our FreshPrompt method leverages a text prompt to (1) introduce contextually relevant and up-to-date information (including answers to relevant questions) from a search engine to a pre-trained LLM, and (2) teach the model to reason over retrieved evidences.
More specifically, given a question , we first use verbatim to query a search engine, in our case Google Search.We scrape the results from Google Search using SerpApi (https://serpapi.com). We retrieve all of the search results, including the answer box, organic results, and other useful information, such as the knowledge graph, questions and answers from crowdsourced QA platforms, and related questions that search users also ask (see Figure 9 in Appendix F). For each of these results, we extract the associated text snippet along with other information, such as source (e.g., Wikipedia), date , title , highlighted words , and then create a list of retrieved evidences . These evidences are then cast into a common format (Figure 3, left) and used to condition the model through in-context learning. To encourage the model to focus on more recent evidences, in line with recent findings (Liu et al., 2023a), we sort the evidences in the prompt from oldest to newest.
To help the model to “understand” the task and the desired output, we provide few-shot demonstrations of input-output exemplars at the beginning of the input prompt. Each demonstration shows the model an example question and a list of retrieved evidences for the question, followed by a chain-of-thought reasoning over the evidences to figure out the most relevant and up-to-date answer (Figure 3, right). Although we include a few exemplars of questions with false premises in the demonstrations, we also experiment with an explicit false premise check in the prompt: “Please check if the question contains a valid premise before answering”. Figure 10 in Appendix G shows a realistic prompt.
2 Experiment setup
We closely follow the setup in Section 3 except in cases where we lack control over the model’s decoding via an API (e.g., Perplexity.AI). Some of the models we evaluate can potentially change over time, which presents a challenge to the reproducibility of our evaluation results; thus, we evaluate all models on the same date of April 26, 2023. In addition to GPT-3.5 and GPT-4, we evaluate Google Search by simply querying Google Search and using the answer in the answer box (if any) or the text snippet of the top-1 search result; Perplexity.AI (PPLX.AI), an answer engine that combines an LLM and a search engine to generate useful responses to users’ queries;https://www.perplexity.ai. At the time of evaluation, PPLX.AI was a combination of GPT-3.5 and Bing Search, and was able to provide both concise and detailed answers. We evaluated its concise answers. and Self-Ask (Press et al., 2022), a method that uses few-shot in-context learning to teach an LLM to decompose each question into simpler sub-questions that are answered via Google Search.We use the few-shot prompt provided by Self-Ask’s authors and apply it to both GPT-3.5 and GPT-4. For simplicity, we evaluate solely the final answer from Self-Ask, disregarding intermediate answers.
FreshPrompt setup: We apply FreshPrompt to both GPT-3.5 and GPT-4 by sequentially incorporating the following retrieved evidences into the input prompt: organic search results, related questions that search users also ask, questions and answers from crowdsourced QA platforms, and the snippets from the knowledge graph and answer box (if available). These evidences are arranged in sequence up to the end of the prompt. Given the models’ context limit, we only keep the top evidences (closer to the end of the prompt) after sorting them based on the corresponding date. Unless otherwise specified, we use for GPT-3.5, and for GPT-4. Additionally, we include question-answer demonstrations at the beginning of the prompt.
3 Results and Discussion
Table 1 presents concrete numbers under Strict (see Appendix H for results under Relaxed). FreshPrompt offers large improvements over the vanilla GPT-3.5 and GPT-4 across the board. GPT-4 + FreshPrompt achieves absolute accuracy improvements of 47% and 31.4% over GPT-4 under Strict and Relaxed, respectively. The reduction in the absolute accuracy gap between Strict and Relaxed (from 17.8% to 2.2%) also suggests that FreshPrompt dramatically diminishes the presence of outdated and hallucinated answers. Unsurprisingly, the most significant improvements for both GPT-3.5 and GPT-4 are on the categories of fast-changing and slow-changing questions, which both concern recent knowledge. That said, questions about old knowledge also benefit from FreshPrompt. For example, GPT-4 + FreshPrompt yields a +30.5% higher accuracy than GPT-4 on questions with valid premises that involve knowledge before 2022 (+9.9% under Relaxed). Additionally, FreshPrompt produces notable gains on false-premise questions (+37.1% and +8.1% respective accuracy improvements under Strict and Relaxed for GPT-4).
FreshPrompt outperforms other search-augmented methods by a large margin:
GPT-4 + FreshPrompt demonstrates superior accuracy across question types, surpassing all other methods by a substantial margin. Its best variant (with 15 retrieved evidences per question) achieves impressive overall accuracies of 77.6% and 79.0% under Strict and Relaxed, respectively. GPT-3.5 + FreshPrompt surpasses PPLX.AI and Self-ask (all performed on top of GPT-3.5) in overall accuracy by +3.8% and +14.4% under Strict. Under Relaxed, however, PPLX.AI achieves a +4.2% higher overall accuracy than GPT-3.5 + FreshPrompt, which is a large part due to its superior accuracy on false-premise questions (58.1% vs. 41.1%). The large accuracy gap of 14.0% between Strict and Relaxed for PPLX.AI suggests that its outputs contain a large amount of hallucination. Overall, all search-engine augmented approaches (Self-ask, PPLX.AI, and FreshPrompt) provide significant gains across question types over vanilla GPT-3.5 and GPT-4. Google Search generally provides better results than both GPT-3.5 and GPT-4, except on questions with false premises, but lags far behind PPLX.AI and GPT-3.5/GPT-4 + FreshPrompt across the board.
The premise check boosts accuracy on false-premise questions but can hurt accuracy on those with valid premises:
As discussed in Section 3.1, OpenAI’s LLMs such as GPT-3.5 and GPT-4 are likely tuned to handle false-premise questions, and this is also true for PPLX.AI. Additionally, we empirically find that several LLMs possess the ability to debunk a false-premise question if explicitly asked, e.g.. “Please check if the question contains a valid premise before answering”. Adding this premise check to GPT-3.5 and GPT-4 yields +23.4% and +6.4% respective accuracy improvement on false-premise questions under Strict (+22.6% and +11.3% under Relaxed). However, this is harmful for GPT-3.5 with regard to other question types, decreasing overall accuracy by 20.8% and 21% under Strict and Relaxed, respectively. This is not a problem for GPT-4, with a slight decrease of 0.6% under Strict and a slight increase of and 1.2% under Relaxed.
Having more relevant and up-to-date evidences at the end of the input context is helpful:
We also analyze how the order of the evidences in the prompt impacts GPT-4’s accuracy. Our results show that using the order returned by Google Search (search order, top search results at the end of the input context) or sorting the evidences by their associated date information (time order, more recent results at the end) generally results in better accuracy compared to using a random order (random order), with up to a +2.2% higher overall accuracy in Strict and Relaxed. Using only the text snippet for each evidence without additional information (such as source, date, etc.) as in GPT-4 + FreshPrompt slightly reduces accuracy, with less than 1% in both settings.
Additional retrieved information beyond the organic search results provides further gains:
Incorporating additional retrieved evidences other than the organic search results, such as the answer box or related questions that search users also ask, is helpful. Removing the answer box decreases GPT-4 + FreshPrompt’s overall accuracy under Strict by 1.4% (1.6% under Relaxed). Removing both the answer box and other relevant information (including related questions) reduces GPT-4 + FreshPrompt’s overall accuracy by 3.2% (3.0% under Relaxed).
Increasing the number of retrieved evidences further improves FreshPrompt:
We explore the effect of the number of retrieved evidences for each question as well as the number of demonstrations by varying these numbers in our experiments with GPT-4. Note that our default setting for GPT-4 + FreshPrompt uses 10 retrieved evidences for each question and 5 demonstrations. Our results suggest that the number of retrieved evidences for each question is the most important ingredient for achieving highest accuracy. Under Strict, increasing this number from 1 to 5, 10, and 15 leads to corresponding overall accuracy improvements of +9.2%, +14.2%, and +16.2%, respectively. This suggests that GPT-4 is able to efficiently handle an increasing number of retrieved evidences (including conflicting answers) and ground its responses into the most factual and up-to-date information. On the other hand, increasing the number of demonstrations from 5 to 15 slightly hurts accuracy in both evaluation settings (1% decrease in overall accuracy under Strict).
Verbose demonstrations improve on complex questions but also increase hallucination:
To evaluate the effect of the writing style of the answer (including the reasoning) in each demonstration, we manually rewrite these answers into a more verbose version (long demonstration answers). Our manual inspection reveals that using more verbose demonstration answers may be helpful when dealing with complex questions but can be more harmful as it provides room for hallucination (a decrease of 2.6% in overall accuracy under Strict).
Related Work
Many prior works study semi-parametric knowledge augmentation in LLMs via additional fine-tuning (Guu et al., 2020; Lewis et al., 2020; Borgeaud et al., 2022; Izacard et al., 2022), while others advocate for knowledge generation instead of retrieval (Yu et al., 2023a; Sun et al., 2023). FreshPrompt aligns with a recent emerging trend in QA applications that augments LLMs’ prompts with knowledge retrieved from search engines for real-time alignment to current and factual information (Nakano et al., 2021; Lazaridou et al., 2022; Menick et al., 2022; Yao et al., 2022; Press et al., 2022; Khattab et al., 2022; Schick et al., 2023; Luo et al., 2023). Similar to our method, Lazaridou et al. (2022) proposed a few-shot in-context learning approach that inserts documents from Google Search into the prompt. We do not compare to this method due to its expensive inference cost, as it chunks retrieved documents into evidence paragraphs and performs inference calls to the LLM to generate answers followed by LLM reranking. In contrast, FreshPrompt only performs a single inference call to the LLM. Self-Ask (Press et al., 2022) also uses few-shot in-context learning to teach an LLM to ask itself follow-up questions before answering the initial question, although it focuses more on decomposition.
Time-sensitive QA:
FreshQA aligns with a growing body of work on benchmarking LLMs’ temporal reasoning capabilities (Chen et al., 2021b; Zhang & Choi, 2021; Liska et al., 2022; Kasai et al., 2022). Chen et al. (2021b) created TimeQA by extracting evolving facts from WikiData along with aligned Wikipedia passages to synthesize 20K timestamped question-answer pairs. Zhang & Choi (2021) constructed SituatedQA by annotating 9K realistic questions from existing open-domain QA datasets with temporal context (i.e., timestamps). StreamingQA (Liska et al., 2022) consists of both LLM-generated and human-written questions (146K total questions) answerable from a corpus of timestamped news articles. Also related is the dynamic RealtimeQA benchmark (Kasai et al., 2022), which evaluates models weekly on a set of around 30 multiple-choice questions about new events extracted from news websites. In contrast, FreshQA contains a fixed set of human written open-ended questions whose answers by nature can change based on new developments in the world and thus offers a complementary generative evaluation of time-sensitive QA.
QA over questionable or counterfactual premises:
Recent work has also introduced QA benchmarks with questionable premises (Yu et al., 2023c; Kim et al., 2023) or counterfactual premises (Yu et al., 2023b). Crepe (Yu et al., 2023c) consists of 8400 Reddit questions (of which 25% questions contain false premises annotated by human workers) split into train/dev/test sets. Kim et al. (2023) constructed (QA)2, an evaluation set of 602 questions based on frequent search engine queries, which are annotated by expert annotators and crowdworkers, and evenly divided between those with and without questionable premises. Consistent with these efforts, we find that current LLMs struggle with handling false premise questions; additionally, several LLMs are able to debunk a false-premise question if explicitly asked to check for the premise’s validity. Similar to above, these benchmarks are complementary and combining them is a promising direction for future work.
Limitations and Future Work
One obvious challenge with FreshQA is the need for regular answer updating by the maintainers; in the interim period between updates, the answers to some questions might become stale. This could be addressed by support from the open-source community (e.g., updates via GitHub pull requests). On the method side, FreshPrompt interfaces with Google Search, and it is unclear how it performs with other search engines for which some types of context (e.g., answer boxes) are not available. Additionally, we only perform one search query per question, and thus our method could be further improved via question decomposition and multiple search queries (Khattab et al., 2022). Since FreshQA consists of relatively simple English language questions, it is also unclear how well FreshPrompt performs in the context of multilingual/cross-lingual QA and long-form QA (Fan et al., 2019). Finally, FreshPrompt relies on in-context learning and thus may underperform approaches that fine-tune the base LLM on new knowledge.
Conclusion
Our work offers a fine-grained and exhaustive evaluation of the capabilities of modern LLMs to adapt to ever-changing world knowledge with and without search engine augmentation. In the process, we develop a new dataset—FreshQA—of 600 questions that test a broad range of reasoning abilities, from the incorporation of fast-changing knowledge to identification of questions with false premises. Our two-mode evaluation also provides a way to measure both correctness and hallucination. Additionally, we propose a simple few-shot in-context learning algorithm called FreshPrompt that incorporates relevant evidences retrieved from Google Search into the prompt of an LLM. FreshPrompt significantly improves performance over competing search engine-augmented approaches on FreshQA, and an ablation reveals that factors such as the number of incorporated evidences and their order impact the correctness of LLM-generated answers. We release FreshQA and commit to updating its answers regularly to facilitate future research.
Acknowledgements
We thank Colin Raffel, Hamed Zamani, and Subhransu Maji for helpful discussion and feedback. We would also like to thank Chengrun Yang, Xinyun Chen for their insightful comments on this manuscript. Finally, we are grateful to the following people for their contributions to creating our FreshQA dataset: Marzena Karpinska, Dustin Tran, Daniel Cer, Sam Fullerton, Elizabeth Clark, Nishant Raj, Xiaoyu Song, Yapei Chang, Yixiao Song, Nader Akoury, Ankita Gupta, Bill Ray, Chau Pham, Wenlong Zhao, Maximilian Mozes, Simeng Sun, Ronan Salz, Kalpesh Krishna, Katherine Thai, Kanishka Misra, Salaheddin Alzu’bi, Erica Cai, Thibault Sellam, Jiao Sun, Dhruv Agarwal, Tessa Masis, Andrew Drozdov, Brian Lester, George Wei, Naveen Jafer Nizar, Shufan Wang, Youngwoo Kim, and Shib Sankar Dasgupta. This project was partially supported by award IIS-2046248 from the National Science Foundation (NSF), as well as NSF’s CloudBank program.
References
Appendix
Appendix A Evaluation protocol
Figure 4 shows specific examples of each evaluation criteria.
Appendix B Inter-rater agreement and automatic evaluation
Two authors independently evaluated a randomly sampled subset of 100 answers across models (including 50 questions with valid premises and 50 questions with false premises) in both modes Relaxed and Strict.
To facilitate future evaluations, we also develop FreshEval, a simple automatic metric that uses few-shot in-context learning to teach an LLM to judge model responses. In each evaluation, the model is conditioned on a given question, a list of valid answers for the question, and a model response, and is then expected to generate a comment on the correctness of the response, followed by a final judgement. At the beginning of each input prompt, we also provide an instruction of the evaluation task, and sample comments and evaluations of the examples in Figure 4 as demonstrations.In our experiments, we found that using separate prompts for Relaxed and Strict evaluations resulted in better performance compared to using a single, combined prompt for both evaluation modes. We also found that additionally incorporating retrieved evidences for the question into the prompt did not improve inter-rater agreement between FreshEval and human raters. See Figure 5 and Figure 6 for FreshEval’s prompts for Relaxed and Strict evaluations, and Figure 7 for FreshEval’s sample output for Strict evaluation.
Table 2 reports the inter-rater agreement between the two human raters, and between FreshEval and each human rater, in terms of exact accuracy. The two human raters had an agreement of 99% for Relaxed and 96% for Strict, while FreshEval achieved an average agreement of 96.5% with human evaluations for Relaxed and 96% for Strict. Overall, the high accuracies demonstrate that our evaluation protocol is reproducible and reliable, and FreshEval can be used in place of human evaluation on FreshQA.
Appendix C Additional experiment setup details for Section 3
To increase reproducibility, we select the most likely token at every decoding timestep (i.e., with a temperature of 0) and generate a maximum number of 256 tokens for all models. Note that the API for some models is non-deterministic by default, even with a temperature of 0. For non-chat models that were not pre-trained with a QA task, we feed them a text prompt of the format: “Q:
For OpenAI models, we use the 2023-03-15-preview API in Azure OpenAI Service. We use the model names text-davinci-003, code-davinci-002, gpt-3.5-turbo, and gpt-4 for GPT-3.5, Codex, ChatGPT, and GPT-4, respectively.
Appendix D Additional experiment results for Section 3
Table 3 and Table 4 show the accuracy of different LLMs on FreshQA under Strict (no hallucination) and Relaxed evaluations, respectively.
Appendix E ChatGPT/GPT-4’s awareness of recent knowledge
Although ChatGPT and GPT-4 were originally trained in 2021, our manual evaluation suggests that they have been exposed to data containing information beyond their knowledge cutoff date in September, 2021. Figure 8 indicates that ChatGPT is aware of the recent Russian invasion of Ukraine on February 24, 2022.
Appendix F Google Search results
Figure 9 shows different types of search results from Google Search for given a query.
Appendix G A realistic prompt for FreshPrompt
Figure 10 displays a realistic prompt for FreshPrompt.
Appendix H Additional experiment results for Section 4
Table 5 presents the accuracy of different search engine-augmented LLMs on FreshQA under Relaxed.