Check Your Facts and Try Again: Improving Large Language Models with External Knowledge and Automated Feedback
Baolin Peng, Michel Galley, Pengcheng He, Hao Cheng, Yujia Xie, Yu Hu, Qiuyuan Huang, Lars Liden, Zhou Yu, Weizhu Chen, Jianfeng Gao
Introduction
Large Language models (LLMs), such as GPT-3 Brown et al. (2020) and ChatGPT, have demonstrated an outstanding ability in generating fluent, coherent, and informative natural language texts. It is commonly understood that the impressive capabilities of these models stem from the abundance of world knowledge encoded therein and models’ ability to generalize from that knowledge. However, the knowledge encoding of LLMs is lossy and the knowledge generalization could lead to “memory distortion.” As a result, these models tend to hallucinate, which can cause damage when deployed for mission-critical tasks. In addition, even with exponentially growing model sizes, LLMs can never encode all information needed for many applications. For example, constant changes in real-world settings cause LLMs to quickly become stale for time-sensitive tasks such as news question answering, and many proprietary datasets are not available for LLM training due to privacy. While there is a growing interest in improving LLMs using external knowledge (e.g., Ghazvininejad et al., 2017; Guu et al., 2020; Zhong et al., 2022; Gao et al., 2019, 2022), almost all the previously proposed methods require finetuning the parameters of a LLM, which can be prohibitively expensive as the size of LLMs grows exponentially. Thus, it is highly desirable to augment a fixed LLM with plug-and-play (PnP) modules for mission-critical tasks.
In this paper, we present LLM-Augmenter to improve LLMs with external knowledge and automated feedback using PnP modules. As illustrated by the example in Figure 1, given a user query (e.g., regarding a 2013 Los Angeles Galaxy player transfer), LLM-Augmenter first retrieves evidence from external knowledge (e.g., Web or task-specific datasets) and, if necessary, further consolidates evidence by linking retrieved raw evidence with related context (e.g., information of the entity “2013 Los Angeles Galaxy”) and performing reasoning to form evidence chains (e.g., table-passage in the figure). Then, LLM-Augmenter queries a fixed LLM (i.e., ChatGPT in our study) using a prompt that contains the consolidated evidence for ChatGPT to generate a candidate response grounded in external knowledge (evidence). LLM-Augmenter then verifies the candidate response e.g., by checking whether it hallucinates evidence. If so, LLM-Augmenter generates a feedback message (e.g., about the team “C.S.D. Municipal”). The message is used to revise the prompt to query ChatGPT again. The process iterates until a candidate response passes the verification and is sent to the user.
In addition to proposing LLM-Augmenter, to be detailed in Section 2, we make the following contributions. We perform an empirical study to validate the effectiveness of LLM-Augmenter using two tasks, information seeking dialog (Section 3) and open-domain Wiki question answering (Wiki QA) (Section 4). The study shows that LLM-Augmenter significantly reduces ChatGPT’s hallucinations without sacrificing the fluency and informativeness of its generated responses. For example, on the dialog task of customer service, human evaluation shows LLM-Augmenter improve ChatGPT by 32.3% in Usefulness (measuring the groundedness or hallucination of model responses) and 12.9% in Humanness (measuring the fluency and informativeness of model responses). The Wiki QA task is extremely challenging to ChatGPT in that answering these questions often requires multi-hop reasoning to piece together information of various modalities scattered across different documents. Our results show that although the closed-book ChatGPT performs poorly and often hallucinates, LLM-Augmenter substantially improves the factuality score of the answers (absolute +10% in F1) by grounding ChatGPT’s responses in consolidated external knowledge and automated feedback.
LLM-Augmenter
The architecture of LLM-Augmenter is illustrated in Figure 2. It consists of a set of PnP modules (i.e., Working Memory, Policy, Action Executor, and Utility) to improve a fixed LLM (e.g., ChatGPT) with external knowledge and automated feedback to mitigate generation problems such as hallucination.
We formulate human-system conversation as a Markov Decision Process (MDP) described by a five-tuple :
is an infinite set of dialog states, which encode information stored in Working Memory, including dialog history, user query, evidence, candidate response;
is a set of actions that Policy picks to execute, including (1) calling Knowledge Consolidator to consolidate evidence from external knowledge and (2) calling Prompt Engine to query the LLM to generate candidate responses;
gives the transition probability of entering a new state after action is taken in state ;
is the external reward received after taking action in state , which is provided by the environment (e.g., users or simulators); and
In what follows, we describe in detail the modules of LLM-Augmenter.
This module tracks the dialog state that captures all essential information in the conversation so far. The state is represented using a six-tuple :
is evidence for , consolidated from external knowledge by Knowledge Consolidator;
is a set of the LLM-generated candidate responses for ;
is a score assessing the utility of each element of , and is a verbalized feedback to guide the LLM to improve its utility — both and are generated by the Utility module (see Section 2.4); and
is the dialog history before .
Note that given user query , LLM-Augmenter can take multiple iterations to revise its response, with each iteration generating a candidate response based on evidence, feedback and utility, before sending the final response to the user, as illustrated in Figure 1.
2 Policy
This module selects the next system action that leads to the best expected reward . These actions include (1) acquiring evidence for from external knowledge, (2) calling the LLM to generate a candidate response, and (3) sending a response to users if it passes the verification by the Utility module.
The policy can be implemented using manually crafted rules, or trained on human-system interactions. In this study, we implement a trainable policy as a neural network model parameterized by . is optimized using REINFORCE (Williams, 1992) to maximize the expected reward as:
We find it effective to implement using a pre-trained model (e.g., T5), which allows us to not only leverage the capacity of the pre-trained model, but also to incorporate additional information through finetuning.
Policy learning typically requires large amounts of human-machine interactions, which can be costly to collect. To address the challenge, policy learning can be done in three stages:
Bootstrapping from a rule-based policy: Domain experts encode task-specific knowledge and business logic into IF-THEN rules. For example, if a product name is mentioned in a user query for customer service, it is wise to always call Knowledge Consolidator to collect information of the product from a product database.
Learning with user simulators: We use a language model to simulate how human users interact with LLM-Augmenter. Any valid response from LLM-Augmenter that passes the evaluation of the Utility module can be used as a training example, allowing LLM-Augmenter to self-improve.
Finally, LLM-Augmenter interacts with human users to further refine its policy.
In addition to Policy, the other trainable modules of LLM-Augmenter (i.e., Knowledge Consolidator and Utility) can also be optimized using the same learning method.
3 Action Executor
This module performs an action selected by the policy. It is composed of two components, Knowledge Consolidator and Prompt Engine.
The Knowledge Consolidator augments LLMs with the capability of grounding their responses on external knowledge to mitigate hallucination when completing tasks, such as answering questions regarding latest news, and booking a table in a restaurant. Following Ma et al. (2022), the Knowledge Consolidator is designed in a modular fashion, consisting of a knowledge retriever, an entity linker and, an evidence chainer.
Specifically, the retriever first generates a set of search queries based on and , and then calls a set of APIs to retrieve raw evidence from various external knowledge sources, such as calling Bing Search APIs to query Web documents including Wiki articles and Reddit messages, and REST APIs to query task-specific databases for restaurant reviews and product specifications.
The retrieved raw evidence is sometimes incomplete and noisy. Thus, the entity linker enriches raw evidence with related context to form evidence graphs, i.e., linking each entity mentioned in raw evidence to its corresponding description based on Wikipedia. Then, the chainer prunes irrelevant evidence from the graphs and forms a shortlist of evidence chains that are most relevant to queries. The consolidated evidence is then sent to Working Memory. Figure 1 shows an example of consolidated evidence for the anchored club “Los Angeles Galaxy”, i.e., two evidence chains corresponding to the transfer players in 2013 season and the former clubs, respectively.
3.2 Prompt Engine
The Prompt Engine generates a prompt to query the LLM to generate a (candidate) response for . The prompt is a text string that consists of task instruction, user query , dialog history , evidence if it is made available by Knowledge Consolidator, and feedback if it is made available by the Utility module. Prompts are task-specific, and details thereof are provided in Appendix A.
4 Utility
Given a candidate response , the Utility module generates utility score and a corresponding feedback using a set of task-specific utility functions.
These utility functions Our experiments are with a single utility function. To allow multiple utility functions, we could learn a linear function mapping the outputs of these multiple functions to a single score using a linear function trained together with the other parameters of the policy. access the alignment of the LLM’s responses with user expectations or specific business requirements. For example, in an information seeking dialog, it is important that all LLM’s responses are preciously grounded in external evidence to avoid generating misleading or inaccurate information. In a restaurant reservation dialog, the LLM responses should be conversational and focused on guiding the user through the reservation process, rather than engaging in off-topic chitchats.
Inspired by Glaese et al. (2022), there can be two distinct types of utility functions:
Model-based utility functions assign preference scores to different dimensions of a response, such as fluency, informativeness and factuality. These functions are trained on pre-collected human preference data or annotated log data.
Rule-based utility functions, implemented using heuristics or programmed functions, measure whether a response complies with a specific rule.
In addition, we have developed a utility function to generate informative and actionable feedback to help revise prompts to allow the LLM to generate better responses. As shown in Figure 1, the utility function generates feedback “but there is no information about the number of international titles.” Such a utility function is a text generation model parameterized by , and can be implemented as a seq2seq or auto-regression language model. It tasks as input user query , evidence , candidate response and dialog history , and generates feedback in text as
Alternatively, LLMs and rule-based natural language generator can be used for feedback generation.
In the next two sections, we present our experiments to validate the effectiveness of LLM-Augmenter in two types of distinct scenarios: (1) information seeking dialog, where the AI agent needs to generate informative and trustworthy responses based on a variety of external sources of knowledge, and (2) Wiki question answering, where the AI agent needs to answer questions by piecing together information of various modalities scattered among multiple Wiki documents.
Information Seeking Dialog
We repurpose the DSTC7 Track 2 task as an evaluation corpus for news conversation. The goal of this task is to generate informative responses that are grounded in external knowledge (i.e., news) and go beyond chitchat. We followed the data crawling process used in DSTC7 Task 2 Galley et al. (2019). We started by selecting Reddit discussion threads that contained URLs in the description, which were crawled from various news-related subreddits during the time period of 2021-2022. We then restricted the URL domain to a curated list of news websites, and extracted the relevant oracle passage by selecting the most appropriate passage for the context based on ROUGE-F1 scores Lin (2004). In order to reduce noisy or irrelevant information, we only kept examples with an F1 score higher than a certain threshold, resulting in a total of 1370 examples for evaluation.
We use DSTC11 Track 5 Kim et al. (2023) as a showcase in a conversational customer service scenario. It expands upon the DSTC9 Track 1 dataset by incorporating subjective knowledge from customer reviews in addition to factual knowledge from FAQs. This allows users to have an engaging and informative conversational experience with the AI system. The dataset evaluates the ability of the AI agent to understand relevant user review posts and FAQs, and generate responses based on both reviews and FAQ snippets. It is collected based on the MultiWOZ 2.1 Eric et al. (2019) dataset and includes users’ knowledge-seeking queries that require the AI agent to use FAQs and user reviews to respond. There are 14768 dialog sessions for training and validation, and the test set is currently unavailable. Therefore, we used the validation set for our evaluations.
2 Experiment Setup
Throughout this work, we focus on using ChatGPT as the backbone black-box LLM. It is straightforward to apply LLM-Augmenter to other LLMs, such as GPT-3 Brown et al. (2020) or PaLM Chowdhery et al. (2022).
For News Chat, Knowledge Consolidator includes a BM25 retriever over web documents linked from Reddit posts. For the Customer Service task, Knowledge Consolidator includes a BM25-based retriever over the knowledge bases of FAQs and Yelp reviews.
Additionally, we also experiment with ground-truth knowledge, referred to as golden knowledge henceforth, which is used by human annotators during data collection, in our oracle experiments.
The prompt templates utilized for News Chat and Customer Service are shown in the appendix in Table 7 and Table 8, respectively.
The goal of this task is to generate responses that are coherent to the context and grounded in external knowledge. To evaluate the degree to which the generated responses are grounded in consolidated evidence, we use the utility score, Knowledge F1 Shuster et al. (2021), to measure the overlap between a prediction and evidence which is either consolidated by Knowledge Consolidator or provided as golden knowledge. Feedback generation is accomplished using a template-based natural language generator.If the KF1 score falls below a certain threshold, the feedback is “The response is inconsistent with the knowledge. Please generate again.” In addition, we use ChatGPT as a utility function, i.e., self-criticism to gather feedback by prompting ChatGPT to evaluate candidate responses and give feedback on how to improve them.
Due to ChatGPT’s current limited bandwidth, we use a rule-based policy for our experiments involving ChatGPT. The prior knowledge about this task inspired us to design a policy that always uses Knowledge Consolidator, evaluates the quality of a candidate response using ChatGPT, and provides feedback to revise the prompt. Additionally, to test the viability of LLM-Augmenter with a trainable policy, we employ offline RL to train the parameters of Policy as Equation 1, where the policy model is based on T5-Base.
We evaluate the performance of LLM-Augmenter on information-seeking dialog tasks using both automatic metrics and human evaluations. Following the literature, we consider commonly used metrics, Knowledge F1 (KF1) and BLEU-4, in grounded conversational response generation and task-oriented dialog. BLEU Papineni et al. (2002) measures the overlap between the model’s output and the ground-truth human response, while KF1 Lian et al. (2019) assesses the overlap with the knowledge that the human used as a reference during dataset collection. Additionally, we include ROUGE-1 Lin (2004) and METEOR Banerjee and Lavie (2005) as these metrics have been found to best correlate with human judgment on the DSTC9 and DSTC11 customer support tasks Kim et al. (2020). We further include BLEURT Sellam et al. (2020), BERTScore Zhang et al. (2019), chrF Popović (2015), which have been shown to be among the best-performing text generation metrics on dialogYeh et al. (2021); Peng et al. (2022). Lastly, we also consider BARTScore as it has been reported to be one of the best model-based metrics Yuan et al. (2021). Given that BARTScore can be interpreted as a log-probability, we report results with its natural exponent (positive scores). Additionally, we perform a turn-level human evaluation to investigate whether responses are (1) useful and (2) human-like. Following the evaluation protocol by Peng et al. (2022), using Amazon Mechanical Turk, we hired master-level workers with lifetime HIT acceptance rate above 95%, and asked them to answer two questions on usefulness (i.e., which response sounds more useful) and humanness (i.e., which speaker sounds more human).
3 Automatic Evaluation Results
Experiment results are shown in Tables 1 and 2. We observe that ChatGPT achieves reasonable performance even in the zero-shot setting. However, with access to golden knowledge, the performance is dramatically improved. This suggests that while LLMs are able to encode a large amount of general knowledge in their parameters, they can still benefit from more specific, targeted knowledge. This is likely because LLMs are designed to handle a wide range of tasks and therefore may not always have access to the most relevant or up-to-date information for a given task. Our experiments show that providing LLMs with task-specific knowledge can significantly mitigate hallucination without sacrificing the fluency and informativeness of model-generated responses. As demonstrated in Tables 1 and 2, LLM-Augmenter mitigates ChatGPT’s hallucination issue on both the news chat and customer service tasks. Specifically, we observe a significant improvement in KF1 scores of approximately 10 and 6 points, respectively, due to the use of evidence retrieved by Knowledge Consolidator.
As listed in Tables 1 and 2, the results of using golden knowledge setting demonstrate that incorporating feedback from the Utility module leads to substantial improvement 3.3 points in KF1 on News Chat and 7.2 on Customer Service, respectively. Similarly, significant improvement can also be observed when using evidence provided by Knowledge Consolidator.
Figure 3 shows the learning curve of LLM-Augmenter on the customer service task. As we do not have an external reward that would require collecting data from real users, we instead define here our reward as the KF1 utility function. This helps demonstrate the effectiveness of LLM-Augmenter in its reinforcement learning (RL) setup. As our experiments are akin to single turn interactions, we did not need to set discount factor , but future work may need to rely on it. We see that LLM-Augmenter’s reward on test data increases as the number of training episodes (dialog sessions) increases, surpassing a random policy after 600 interactions and ultimately reaching a KF1 score of approximately 37.5. Through these interactions, LLM-Augmenter is able to learn to effectively select the next system action to maximize the reward, which helps our system reduce hallucinations while generating fluent and informative responses.
4 Human Evaluation Results
We compare ChatGPT with and without LLM-Augmenter. A total of 948 randomly selected examples from the customer service dataset are used for human evaluation. The evaluation results are converted from a 5-point Likert-like scale to a win/tie/loss scale for reporting, as shown in Table 3. We observe a strong preference for LLM-Augmenter over ChatGPT alone in terms of both Usefulness and Humanness. The result is consistent with the automatic evaluation result, discussed earlier.
5 Ablation Study
We conduct ablation experiments to evaluate the effect of various policies on the utilization of the knowledge consolidator. Figure 4 shows the performance of three different variants of the policy: 1) no-knowledge consolidator, in which the knowledge consolidator is not used, 2) Self-ask, in which the knowledge consolidator is only utilized when the LM suggests the use of external knowledge by prompting it whether to use, and (3) Always-use, in which the knowledge consolidator is always provided to the LM. Our results indicate that Self-ask policy achieves a significantly better KF1 score than the No-knowledge consolidator policy, with the ChatGPT model unable to answer user queries and suggesting knowledge consolidator access for 24% of examples. However, the Always-use policy, while achieving the best KF1 score, also incurred additional overhead in terms of knowledge consolidator access. These observations suggest that a trainable policy model should be employed to learn when to use external knowledge.
In addition, the evaluation results on the impact of different types of feedback for LLM-Augmenter are listed in Table 4. We observe that self-criticism feedback enhances response quality make it more knowledge-grounded. Although its performance is comparable to that of rule-based feedback, it provides more detailed suggestions. We speculate that self-criticism will be more helpful for complex tasks. Some examples can be found in 6.
To understand the impact of utility functions and feedback-augmented prompting on the performance of LLM-Augmenter, we conduct an analysis by turning each component on and off. Figure 5 illustrates the results of each variant. We observe that the combination of using utility functions and feedback-augmented prompting, i.e., 2, achieves the best performance. In addition, always providing feedback (as shown in 4) also enhances the performance, although it requires additional model prompting. 5 represents prompting ChatGPT twice and re-ranking the response based on the utility functions, which results in a slightly higher KF1, but performs significantly worse than 2. These findings suggest that incorporating both utility functions and feedback is a more effective method for improving the alignment of LLMs.
Wiki QA
Instead of conversational evaluations, we focus on stress tests on ChatGPT here using open-domain question answering. As ChatGPT and other LLMs are mostly trained using abundant text from single web pages, we hypothesize that answering multi-hop questions involving scattered information across different pages/modalities can better serve the purpose. Due to this, closed-book LLMs are more likely to hallucinate. Moreover, the complex step-by-step reasoning can even be challenging for existing search systems to gather all necessary support evidence in one-shot. Thus, more advanced knowledge consolidation techniques are essential to elicit LLMs for proper grounding. Lastly, different from conversational tasks where long-form responses are desirable, we mainly consider questions with concise short-form answers, i.e., there exists a significant style shift in responses. To align ChatGPT to this new scenario with distinct characteristics, extra instructions are needed.
The OTT-QA dataset is an open-domain question answering benchmark that considers multi-step joint reasoning over both tabular and textual information. It consists of around 40K instances built upon Wikipedia, including 400K tables and 6M passages as the knowledge source. Solving the questions in OTT-QA requires diverse reasoning skills and can be divided into three categories: single-hop questions (13%), two-hop questions (57%), and multi-hop questions (30%). In this paper, we denote the dataset as Wiki QA.
2 Experiment Setups
In the following, we describe the experimental setup for Wiki QA. Unless specified otherwise, the setups are identical to those used in Section 3.
Here, the Knowledge Consolidator uses Wikipedia passages and tables as the knowledge source. Instead of using BM25 as done for dialog tasks, we resort to a dense model, DPR Karpukhin et al. (2020), as the backbone retriever. For DPR, both question and passage/table inputs are represented by the corresponding special token [CLS] embeddings from their respective encoders, and retrieval is simply done via maximum inner product search in the vector space. Given a question, we use DPR to obtain the initial set of evidence, which includes tables and passages. As most WikiQA questions require reasoning hops across different pieces of information (e.g., hopping from the album table to its entry artist page in Figure 1), we contend that directly feeding this raw evidence set to Working Memory is insufficient for prompting LLMs. Thus, we further use additional intermediary modules, i.e., linker and chainer, from CORE Ma et al. (2022) to consolidate the raw evidence, including connecting relevant documents, reranking evidence, and splicing them into evidence chains. We refer to Ma et al. (2022) for more details.
The prompt templates utilized for Wiki QA is shown in the appendix in Table 9.
Here, as a response to a given question is deemed to leverage information from the consolidated knowledge, we use recall as the utility score, i.e., preferring responses with higher token overlap with the corresponding evidence set. Similar to Section 3, we again consider a template-based natural language generator for giving feedback to ChatGPT.
As WikiQA mainly concerns short-form answers, we evaluate the generated responses using the token-level precision, recall and F1 scores against the annotated answers.
3 Results
Table 5 presents the evaluation results on Wiki QA. As expected, the closed-book model alone performs very poorly. Based on our manual inspections, we find that most error cases are hallucinated answers and ChatGPT abstains from answering for cases. We observe that incorporating knowledge obtained from either DPR or CORE significantly improves the F1 score. The substantial improvements observed over the closed-book ChatGPT model indicate the importance of enhancing LLMs with external knowledge. Compared with raw evidence from DPR (row 2), we observe that consolidated evidence from our proposed Knowledge Consolidator with CORE (row 3) is more useful to the frozen ChatGPT model, achieving more pronounced improvements across the board. This suggests that it is crucial to consolidate knowledge for eliciting black-box LLMs to perform grounded reasoning. Lastly, consistent with the observations for news chat and customer service scenarios in Section 3, augmenting ChatGPT with automated feedback further improves alignments (adapting ChatGPT to perform multi-step grounded reasoning), leading to a substantial increase in recall and F1 scores.
Compared with the state-of-the-art fine-tuned model Ma et al. (2022) using top-50 consolidated evidence, there still remains a noticeable gap in performance. Besides a lower answer recall of the consolidated evidence, we attribute it to extra alignments required for ChatGPT to respond in a more concise way and conduct faithful step-by-step reasoning. Therefore, there is ample room for future explorations on elicitive prompting to achieve further improvements.
Related Work
Numerous LLMs for text generation Radford et al. (2018) have been proposed over the years, including very competitive ones such as GPT-3 Brown et al. (2020); Ouyang et al. (2022), OPT Zhang et al. (2022), GPT-j Wang and Komatsuzaki (2021), and ChatGPT. However, most of them do not naturally incorporate external knowledge. To address this limitation, various works augment LLMs with knowledge consisting of e.g., personalized recommendations Ghazvininejad et al. (2017), Wikipedia article and web search Dinan et al. (2018); Shuster et al. (2022), structured and unstructured knowledge of task-oriented dialog Peng et al. (2022). Recent advances have focused on jointly finetuning the retriever and generation components of retrieval-augmented text generation systems Lewis et al. (2020); Zhang et al. (2021), but these methods are not applicable to black-box LLMs.
More recent work attempts to combine black-box LLMs with external knowledge, such as incorporating external knowledge into prompts Madaan et al. (2022); Lazaridou et al. (2022), making GPT-3 more faithful He et al. (2022), and combining web knowledge with GPT-3 Nakano et al. (2021). In very recent works related to ours, Shi et al. (2023) tune the ranker of a black-box LLM. Schick et al. (2023) tune black-box LLMs’ access to different APIs and show improvement on a variety of understanding and reasoning tasks. We consider these works to complementary to ours, as we assume our set of APIs to be given and fixed, and we instead focus more on when and what APIs to request, interactive feedback with the LLM, and developing a self-learning ability through utility functions.
Limitations and Future Directions
A main limitation of this work is that interactive feedback with a computationally expensive model such as ChatGPT can significantly slow down the user experience, as ChatGPT is often queried twice for a single response. However, we think this can translate into more choice for the user. For example, the initial ChatGPT response can be shown to the user as it is being decoded, and the user could then be informed that a more accurate response is available (depending on the utility function). Then, an impatient user can decide to ignore this option, while a user more mindful of response accuracy may decide to see the improved ChatGPT response. In task-oriented and high-stakes scenarios, we believe many users would prefer the slower but more accurate option.
The main results of the paper are with a policy designed manually, as due to the current high-demand for ChatGPT and its limited bandwidth. As reinforcement learning can be quite sample inefficient, we trained our policy using an LLM (T5-Base) we could easily query, and these RL experiments demonstrate the effectiveness of LLM-Augmenter. As ChatGPT becomes more available, we plan to update the paper with RL experiments involving ChatGPT. The current version of the paper does not include human evaluation, as the goal with our current utility function (KF1) shown we can make ChatGPT more grounded and our experiments suggest the responses of our best system are better at capturing the words of the (gold) knowledge. As we move to towards much utility functions such as safety, it will be important to add more fine-grained analyzes of the responses, and we will add human evaluation. In future work, we also plan to leverage interactions with real users and user feedbacks to train LLM-Augmenter.
Conclusions
We introduced LLM-Augmenter, a framework for augmenting black-box LLMs (e.g., ChatGPT) with external knowledge and automated feedback. The external knowledge provided as part of the LLM prompts helps generate more responses that are more grounded into external knowledge relevant to the current conversation. The automated feedback elicits the “follow-up correction” abilities of models such as ChatGPT and InstructGPT in order to produce revised responses that rank higher according to some given utility functions (e.g., groundedness as measured by KF1). These various components are integrated together as part of an RL framework, which we optimize end-to-end using policy gradient. End-to-end experiments with T5 show the effectiveness of LLM-Augmenter, while experiments on ChatGPT show significant increases both in terms of KF1 and a host of text generation metrics.
Ethics Statement
It is widely understood that large language models have the potential to generate harmful, offensive, and inappropriate content Bender et al. (2021); Bommasani et al. (2021); Weidinger et al. (2021). This paper is an attempt to address a major harm of LLMs, namely factual integrity. This paper does not address the problem of offensive content generation, but future work on LLM-Augmenter could help mitigate such harm via, e.g., offensiveness-related utility functions.
As with other knowledge-augmented text generation applications, we cannot rule out that external sources could compromise the factuality of generated text. It is, therefore, important to encourage users to check the relevance of external sources that supplement the generated text.
Acknowledgements
We thank Saleema Amershi, Ahmed Awadallah, Nguyen Bach, Paul Bennett, Chris Brockett, Weixin Cai, Dhivya Eswaran, Adam Fourney, Hsiao-Wuen Hon, Chunyuan Li, Ricky Loynd, Hoifung Poon, Corby Rosset, Bin Yu, Sheng Zhang, and members of the Microsoft Research Deep Learning group for valuable discussions and comments.
References
Appendix A Appendix
Table 6 provides sample responses contrasting ChatGPT and LLM-Augmenter. First, we can see that ChatGPT fails to provide a response related to specific knowledge related to the user, e.g., a local Indian restaurant. In the second part of the table, we show LLM-Augmenter’s Working Memory, which highlights the richer information retrieved from external knowledge to help the underling LLM (i.e., ChatGPT as well) generate more contentful responses. The first LLM response received by LLM-Augmenter is unfortunately not satisfactory, as the quality and specificity of LLM generation can be unpredictable. In this case, the Utility module has determined that the first response did not meet its criteria (i.e., KF1 above a given threshold), and issues a feedback to the LLM module (i.e., “response is inconsistent with the knowledge”). The second response received by LLM-Augmenter is much more satisfactory according to the utility function, and therefore sent to the user.