R-Tuning: Instructing Large Language Models to Say `I Don't Know'

Hanning Zhang, Shizhe Diao, Yong Lin, Yi R. Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, Tong Zhang

Introduction

Large language models (LLMs) have demonstrated remarkable performance across numerous tasks; however, they are also plagued by various issues, such as the propensity of large models to fabricate non-existent facts, a phenomenon commonly referred to as hallucination (Maynez et al., 2020a). Towards mitigating the hallucination, current mainstream approaches include retrieval-based methods (Peng et al., 2023; Li et al., 2023b; Luo et al., 2023), verification-based methods (Manakul et al., 2023; Elaraby et al., 2023; Cohen et al., 2023; Du et al., 2023; Gou et al., 2023), and so forth.

In this paper, we first identify the cause of the hallucination, attributing it to the significant gap existing between the knowledge of human-labeled instruction tuning datasets and the parametric knowledge of LLMs. In the process of developing a large model, previous studies (Min et al., 2022; Wang et al., 2023; Zhou et al., 2023) demonstrate that almost all knowledge is acquired in the pre-training stage, while instruction tuning teaches formatting and chain-of-thought prompting guides knowledge elicitation. Consider Figure 1 as an example. During pre-training, models internalize a large volume of factual knowledge, compressing it within their parameters and the fine-tuning process may carry data that is out of the parametric knowledge. However, traditional fine-tuning methods force the models to complete the sentence. Even when faced with questions beyond their knowledge boundary, they venture to provide an answer. Training a model exclusively on correct answers inadvertently teaches it to guess rather than admit its ignorance. Consequently, if we never train the model to articulate "I don’t know" as a response, it remains unequipped to do so when confronted with unknowns. Addressing this challenge, we assert that enabling a model to astutely respond based on its own knowledge limit is of paramount importance. This example illustrates our motivation to tune our model on the intersection of parametric knowledge and the instruction tuning data, leading to a model refusing to answer unknown questions.

In light of this, we propose a novel instruction tuning method, Refusal-Aware Instruction Tuning (R-Tuning). R-Tuning aims to endow the model with refusal-aware answering ability by recognizing when they should — and shouldn’t — claim knowledge. Specifically, R-Tuning introduces two steps: (1) measure the knowledge gap between parametric knowledge and the instruction tuning data, and identify uncertain questions. By inferring the model on the training data once and comparing the prediction and label, the instruction tuning data is split into uncertain data D0D_{0} and certain data D1D_{1}. (2) construct the refusal-aware data by padding the uncertainty expression after the label words, and then finetune the model on the refusal-aware data.

We conduct two types of experiments: single-task and multi-task, with seven datasets. In the single-task experiments, R-Tuning demonstrates the ability to refuse to answer uncertain questions and improve the accuracy of the willingly answered questions. In the multi-task setting, our method not only demonstrates the advantages of multi-task learning on in-domain datasets but also exhibits superior generalization performance on out-of-domain datasets. This verifies that refusal-aware answering is a kind of meta ability, which is not dependent on a specific task and could benefit from multi-task training and joint inference. With more downstream tasks, R-Tuning could abstract and learn such meta ability better.

One way to interpret our method is that it involves learning the uncertainty of the training data as part of instruction tuning. Further analysis surprisingly shows that learning uncertainty during training and then using it to filter and respond to questions yields better results than directly applying uncertainty filtering on test data. This finding suggests that learning uncertainty improves the model’s training in both estimating uncertainty and answering questions. This discovery highlights the advantages of incorporating uncertainty learning into large model training, both in reducing computational overhead during testing and in improving overall model accuracy.

We investigate the knowledge gap present between the instruction tuning data and the parametric knowledge and attribute the hallucination issue to forcing the model to complete answers with traditional instruction tuning.

To address this issue, we propose a novel instruction tuning approach, R-Tuning, that distinguishes instruction tuning data based on the model’s own knowledge. R-Tuning constructs a refusal-aware dataset and then tune the model to refrain from responding to questions beyond its parametric knowledge.

Experimental results demonstrate the effectiveness and generalization abilities of R-Tuning. We find that the model’s learned refusal ability functions as a meta-skill, being task-agnostic and enhanced through multi-task training.

Refusal-Aware Instruction Tuning

In this section, we first introduce the refusal-aware instruction tuning method (R-Tuning), the core idea of which is divided into two steps: the first step involves identifying and recognizing the uncertain data within the instruction tuning dataset, which are beyond the parametric knowledge boundary of the original model. The second step is to construct certain and uncertain data. Then, we will detail the instruction tuning and inference extraction process. An illustration of R-Tuning is shown in Figure 2.

The first step of R-Tuning is to measure the model’s knowledge gap between the parametric knowledge of LLMs and the instruction tuning data. It asks for the model’s prediction when given a question and applies certain metrics to determine when the model does know. Given a training dataset D={(q1,a1),(q2,a2),...,(qn,an)}D=\{(q_{1},a_{1}),(q_{2},a_{2}),...,(q_{n},a_{n})\} consisting of nn question-answer pairs, we introduce a supervised identification strategy. We first apply the pre-trained model MM to answer all the questions in DD and split the questions into two sets based on the comparison between the prediction and label. If the model’s prediction matches the label, the question is assigned to the certain set D1D_{1}, and otherwise, it belongs to the uncertain set D0D_{0}. As shown in Figure 2, in the left part, because the prediction (Beijing) matches the ground-truth label (Beijing), it belongs to certain data D1D_{1}, demonstrating that the model’s parametric knowledge possesses the capability to answer this question. On the contrary, in the right part, the mismatch between the prediction and the ground-truth label results in this question being categorized into uncertain data D0D_{0}. Finally, the training dataset would be split into two sets (i.e., D0D_{0} and D1D_{1}) with the recognition of the knowledge gap between parametric knowledge and the knowledge that training questions require.

In addition to this supervised strategy requiring a ground-truth label, we also explore unsupervised one, which is discussed in the analysis (Section 5.2) and is found to be effective.

2 Refusal-Aware Data Construction

The refusal-aware data is further constructed by incorporating a prompt template. We introduce a padding method, which keeps the original labels while appending the uncertainty expression at the end. The template is

The certain dataset D1D_{1} is constructed by appending "I am sure" after the template, while the uncertain dataset D0D_{0} is constructed by appending "I am unsure" after the template. The prompt we are using is Are you sure you accurately answered the question based on your internal knowledge? As shown in Figure 2, by appending certain and uncertain expressions, R-Tuning teaches the model to express uncertainty toward questions. This template provides all label knowledge to the model while instructing them to express uncertainty at the same time. On the contrary, we can also directly replace the label word with uncertainty expressions for uncertain questions. We call this strategy the replacement method and investigate the effectiveness of it in Section 5.3.

3 Training and Inference

With the refusal-aware dataset, we then apply the standard procedures of fine-tuning a language model. The model takes a sequence t1,t2,…,tTt_{1},t_{2},\ldots,t_{T} consisting of the question and answer, and predicts the answer part based on the question. The training objective is the standard cross-entropy loss L\mathcal{L} which can be defined as:

Here, P(ti∣t1,t2,…,ti−1)P(t_{i}|t_{1},t_{2},\ldots,t_{i-1}) is the probability of the ithi^{th} token tit_{i} given the preceding tokens t1,t2,…,ti−1t_{1},t_{2},\ldots,t_{i-1}, as predicted by the language model. Note that we calculate the loss solely for the answer part, while excluding the loss attributed to the question part.

During the inference, we first fit the input question into the template 1 and the model will output its answer. Then the designed prompt template Are you sure you accurately answered the question based on your internal knowledge? I am will be appended to the question and answer. Based on this prompt, the model can output its confidence about the previous context.

Experimental Settings

In this section, we first provide an overview of the benchmark datasets and the corresponding evaluation settings. Then the baseline models and the implementation details are presented in the following subsections, respectively.

Given the diverse data formats across different tasks, we first unify the downstream data into two formats:

Question-Answering: Given a question, the model directly predicts its answer. We include ParaRel (Elazar et al., 2021), HotpotQA (Yang et al., 2018), SelfAware (Yin et al., 2023), HaluEval (Li et al., 2023a), FalseQA (Hu et al., 2023), and NEC in our experiments.

Multiple-Choice: Given a question with several choices, the model chooses one option. We include MMLU (Hendrycks et al., 2021), WiCE (Kamoi et al., 2023), and FEVER (Thorne et al., 2018) in our experiments.

For question-answering tasks, to compare the answer generated by our model with the ground-truth answer, we examine whether the first few output tokens contain the ground-truth answer. We don’t adopt exact matching (EM) as the generation is not strictly controllable. For multiple-choice questions, we restrict the model to generate one token and select the choice with maximum probability among the candidate choices by argmax⁡x∈Clogits(x),\operatorname{argmax}_{x\in C}logits(x), where CC is the set of candidate choices. Considering the huge size of HotpotQA and FEVER, we randomly sample 10K training data from them, respectively. More details about the original datasets are shown in Appendix A.1 and Table 8. In Figure 6, we present the distribution of constructed refusal-aware data D0D_{0} and D1D_{1}.

Single-task: The single-task experiments verify the effectiveness of learning on individual tasks. We conduct experiments on ParaRel and MMLU datasets, respectively. We manually split the datasets into the training set, in-domain test set, and out-of-domain test set.

Multi-task: The multi-task experiments aim to evaluate the model’s generalization performance. We choose five datasets - ParaRel, MMLU, WiCE, HotpotQA, and FEVER, and mix them to construct a new training dataset. As for testing, we evaluate the performance on their corresponding test set (in-domain) and an unseen test set (i.e., HaluEval) (out-of-domain).

2 Baselines

We consider three baseline models as follows:

Pretrain-T: Evaluate the performance of original pre-trained checkpoints on the entire test set.

Pretrain-W: To verify the effectiveness of willingly answered questions, we evaluate the performance of the original pre-trained checkpoints on the test set that our fine-tuned models are willing to answer. Intuitively, if the willingly answered questions are within the base model’s knowledge, this baseline should perform well.

Vanilla: Fine-tune the model on DD with all questions and ground-truth labels. This is the traditional instruction tuning method.

3 Evaluation

For models that could only output either the answer or an unknown expression, we evaluate the questions that our model is willing to answer. The accuracy is calculated as follows:

For R-Tuning, because it could output both the question’s answer and the uncertainty, we first prompt the model to provide an answer and then prompt it to provide its uncertainty. Then we can evaluate the precision-recall tradeoff based on the uncertainty and prediction performance. We introduce the Average Precision (AP) score, which measures the precision in identifying and ranking relevant predictions. AP score originates from the object detection field (Everingham et al., 2010) by ranking the prediction results by confidence from high to low and calculating the precision at each threshold. The AP score is the average of these precisions, which is calculated as follows:

where nn is the number of data, kk is the number of data we select for the current threshold. PP and RR denote precision and recall, which are defined as

An ideal model predicts the correct answers with high confidence and the hallucinated wrong answers with relatively low confidence, leading to a high AP score. On the other hand, the AP score is low if the model predicts every answer with high confidence, as the precision at every threshold will not be high and the average will be relatively low. In our experiments, the confidence of R-Tuning is calculated as the weighted average of prediction confidence and uncertainty confidence.

4 Implementation

We choose OpenLLaMA-3B (Geng and Liu, 2023), LLaMA-7B, and LLaMA-13B (Touvron et al., 2023) as the base models in our experiments. We use LMFlowhttps://github.com/OptimalScale/LMFlow (Diao et al., 2023a) to conduct instruction tuning, setting epoch to 11 and learning rate to 2e−52e^{-5}. The batch size is 44 and the temperature is . All the experiments are implemented on Nvidia A100-40GB GPUs.

Experimental Results

In the main experiments, we conduct single-task experiments to verify the model’s refusal-aware answering ability and multi-task experiments to investigate the generalization of refusal ability.

We first conduct single-task experiments on ParaRel and MMLU datasets. The results are shown in Figure 3 and Table 1. Firstly, we observe that R-Tuning significantly outperforms other baselines by a large margin in terms of accuracy on the questions it is willing to answer, compared with others that simply answer all the questions. The results first demonstrate the effectiveness of the refusal-aware answering ability. We also conclude that R-Tuning answers more questions within its parametric knowledge during pre-training, which is reflected by the high accuracy of Pretrain-W. Overall, it is observed from Table 1 that R-Tuning outperforms Vanilla in terms of the AP score, demonstrating the benefits of only answering the questions that align with the model’s parametric knowledge with high confidence. In addition, we find that larger models achieve more improvement compared with over baseline as the gap of the AP score becomes larger, indicating good scalability of R-Tuning. For example, R-Tuning of 13B performs significantly better than Vanilla on out-of-domain ParaRel and in-domain MMLU. In addition, the AP score of R-Tuning grows steadily when the model size becomes larger, while the AP score of Vanilla drops in ParaRel (OOD) and MMLU (ID). This comparison shows that Vanilla may suffer from confidence miscalibration problems while R-Tuning is more well-calibrated in terms of confidence. By combining the prediction confidence and certainty confidence to evaluate the output, R-Tuning is more reliable when making predictions.

2 Multi-task Experiments

The results of multi-task experiments are shown in Figure 4. Overall, R-Tuning consistently outperforms all baseline models in terms of the AP score on both ID and OOD tasks, demonstrating its superiority by introducing the refusal-aware dataset. A higher AP score signifies that the R-Tuning has successfully ranked correct answers higher than incorrect answers, demonstrating its effectiveness in accurately identifying the desired predictions. Especially, on the unseen dataset HaluEval-QA, R-Tuning also achieves a higher AP score and demonstrates its ability to express certainty to questions from other distributions, and such ability can be generalized well. The experiments on multi-task datasets tell us that the refusal is a kind of meta-skill of models and could be enhanced by several different datasets. We provide the detailed AP scores and curves for different datasets and model sizes in Table 12 and Figure 8 in Appendix A.6.

In summary, R-Tuning reduces hallucinations by disregarding inquiries outside of the model’s knowledge domain. Meanwhile, R-Tuning performs well with inquiries that are aligned with the model’s parameterized knowledge. The better AP score demonstrates a good trade-off between precision and recall and the performance on multi-task experiments demonstrates the generalization potential of refusal-aware answering ability.

Analysis

In this section, we first provide an interpretation from the uncertainty perspective for R-Tuning and then investigate two variants: R-Tuning with unsupervised identification strategy (R-Tuning-U) and R-Tuning with label replacement (R-Tuning-R). In addition, we verify the refusal ability on unanswerable questions, which should not receive answers from the model. Finally, we investigate the perplexity of refused questions and the entropy of answers to get a deep understanding of how R-Tuning works.

One perspective on interpreting our method is that R-Tuning of selecting and learning through uncertainty fundamentally involves learning the uncertainty of the training data. A more direct baseline is to perform vanilla fine-tuning and then use uncertainty selection on the test dataset to respond, a method we refer to as Vanilla-C. Vanilla-C prompts the model to answer kk times and choose the majority as the answer. The uncertainty is proportional to the distinct answers. In our experiment, we set k=10k=10 for Vanilla-C and the confidence is calculated by:

where nn is the number of distinct answers generated, and fif_{i} is the number of occurrences of ii-th answer. We calculate the AP scores and compare them with R-Tuning in Table 3. Surprisingly, we find that learning uncertainty and then filtering questions based on this uncertainty to provide answers yields better results than directly filtering and answering questions using uncertainty on the test dataset. In other words, differentiating instruction tuning data based on uncertainty while learning both the correct answers and uncertainty not only enables the learning of uncertainty expressions but also, remarkably, improves the accuracy of question-answering. This is an unexpected but intriguing phenomenon. Learning uncertainty from training data should not be as accurate as using uncertainty estimations directly from the test data. One possible explanation is that for a Transformer model, to accurately predict the last token, the hidden states are adjusted during training. These changes in hidden states might help in better answering easier questions. A potential hypothesis is this: predicting uncertainty embeds information about confidence into the hidden representation. This aids in generating more confident hidden states when answering easier questions. This discovery reveals the benefits of learning the uncertainty of large models. It not only avoids the extensive overhead of repeatedly calculating uncertainty during testing but also improves training quality by learning uncertainty, thereby enhancing the accuracy of uncertainty estimation We plan to conduct further experiments to delve deeper into this phenomenon.

2 Unsupervised Identification

During the refusal-aware data identification process, we apply a supervised way to identify unknown questions by comparing the predictions and labels. In this section, we introduce an unsupervised identification method, R-Tuning-U, where the refused questions are determined by the uncertainty of the model. Specifically, R-Tuning-U queries the model MM kk times and calculates the uncertainty uu across kk predictions, which is calculated by the entropy based on kk answers as follows:

where p(aj∣q)p(a_{j}|q) is the frequency of a certain predicted answer aja_{j} given a question qq.

Then the questions could be ranked according to the uncertainty score uu. For the 50%50\% most uncertain questions, we append the ground truth label and uncertain expression (i.e., uncertain set D0D_{0}), while the remaining (i.e., certain set D1D_{1}) are appended with the ground truth answers with certain expressions. We set the temperature to 0.70.7 and k=10k=10 in our experiments. We compare the performance with the R-Tuning on the ParaRel dataset, and the results are shown in Table 4. It is observed that R-Tuning-U generally achieves a higher AP score, which reveals the feasibility of constructing refusal-aware training data by uncertainty. Comparing the output of the pre-trained model with the ground-truth answer is not the only way to evaluate its parametric knowledge. Uncertainty can also be an indicator of whether the pre-trained model is familiar with the knowledge. An advantage of R-Tuning-U is that it does not require the labels of uncertain questions. If combined with the replacement method, it reduces dependence on data annotation. We leave this for future work.

3 Label Replacement

In the main experiments, we adopt the padding method for data construction. In addition to padding, we can directly replace the label words with uncertainty expressions for uncertain questions and keep the original label words for certain questions, which is called the replacement strategy. For example, the certain part of the training questions D1D_{1} is constructed as follows:

while the uncertain dataset D0D_{0} is constructed as follows:

There are many different ways for the uncertainty expression. To increase the diversity, we take the 16 expressions of uncertainty text from Yin et al. (2023). These 16 expressions are listed in the Appendix Section A.3.

We conduct experiments with R-Tuning-R on ParaRel and MMLU datasets by comparing it with vanilla fine-tuning strategy and the original pre-trained models. The results are shown in Figure 5. Firstly, on both in-domain and out-of-domain test sets, the accuracy of R-Tuning-R is higher than Pretrain-T, which benefits from only answering certain questions. More detailed results with answer rate are reported in Table 9, where we find the model is able to refuse a certain amount of questions. Then, R-Tuning-R outperforms Vanilla with a significantly higher accuracy on its willingly answered questions, which demonstrates the effectiveness of our method. It is promising as R-Tuning-R is trained with fewer ground-truth labels, while Vanilla is trained on all labels of the full training data. Generally, larger models possess more powerful refusal abilities. In Figure 5, we observe that on the willingly answered questions, larger models achieve a higher accuracy. In addition, the high accuracy of Pretrain-W reveals that those selected questions are within parametric knowledge of the pre-trained model. In summary, compared with vanilla fine-tuning, R-Tuning-R provides the model with the refusal ability to refuse unknown questions, which eventually improves the accuracy and prevents them from making hallucinated answers. Table 2 shows the case studies of how R-Tuning-R works. There are significant differences when they encounter questions out of their knowledge. The Vanilla model is proactive in making up an answer, which is a hallucination and makes no sense. However, R-Tuning-R refuses them explicitly with keywords do not know, not known, and impossible. The ability of R-Tuning-R to refuse unknown questions results in fewer hallucinations.

Despite this refusal ability, there are two issues with R-Tuning-R: (1) the replacement method throws away valuable labels which could be leveraged for training. (2) R-Tuning could either only output the answer or only output the certainty, but cannot respond to both, leading to difficulties in considering the precision and recall simultaneously. To leverage all ground-truth labels during the tuning process, and instruct models to predict answers and express uncertainty at the same time, we employ the padding strategy in our main approach, where every question is appended with the ground-truth label and the uncertainty expression, indicating whether the model is confident or not.

4 Unanswerable Questions

In addition to the open-ended question-answering dataset where all the questions are answerable, we also test the performance of R-Tuning on several refusal benchmarks containing unanswerable questions. These questions either contradict common sense or make up some concepts, and should not receive answers from the model. We verify R-Tuning and R-Tuning-R on such datasets, and the results are shown in Table 5. For baseline models, we provide explicitly in the prompt that they could refuse to answer the questions. We observe that R-Tuning and R-Tuning-R refuse nearly all these unanswerable questions, which meet our expectations, while other baselines answer most of the questions even though they are told to refuse. In conclusion, both the R-Tuning and R-Tuning-R possess the ability to refuse questions that contradict common sense or out of their parametric knowledge.

5 Perplexity of Datasets

Perplexity measures how well the language model predicts a given text. Lower perplexity means better prediction and understanding of the text. According to the refusal-aware data identification, we split the training data into two sets: D0D_{0} (uncertain questions) and D1D_{1} (certain questions). To uncover why the pre-trained model responds to them differently, we calculate the average perplexity on these two datasets with the pre-trained models. The perplexity is calculated as follows:

where XX denotes a sentence consisting of tokens and X=(x1,x2,…,xt)X=(x_{1},x_{2},\ldots,x_{t}). Specifically, we calculate the perplexity of the training questions to estimate the pre-trained model’s understanding of them. The results are shown in Table 6. We observe that D1D_{1} has a lower perplexity, demonstrating that the pre-trained model is more familiar with the questions and is likely to provide the correct answer. For D0D_{0} data, its higher perplexity shows that these questions are not familiar to the model and out of the model’s knowledge, and this is the reason why the model tends to hallucinate text instead of providing the correct answers. We also observe that larger models have a lower perplexity and randomness on the questions, which is why larger models generally perform better on various tasks.

By instructing our model to express uncertainty toward relatively random questions in terms of perplexity, the model develops a better understanding of uncertainty and ambiguity and learns the ability to recognize when it does not know. This ability is crucial in situations where simply providing a definite answer may be inappropriate or even harmful. On the other hand, since our model is also trained with data with certain expressions, it becomes more proficient at handling less random questions, and answering them with confidence and certainty. Overall, R-Tuning improves the model’s ability to adapt to different levels of question randomness.

To verify the pre-trained model is less familiar with the uncertain questions while more confident with certain questions, we also plot the confidence distribution on certain questions and uncertain questions, shown in Figure 7 in Appendix A.5. It is observed that a larger percentage of certain questions occupies the high confidence intervals, which means when the model provides correct answers, it generally shows larger confidence.

6 Entropy of Answers

In addition to evaluating the difference between certain and uncertain questions with pre-trained models, we further leverage GPT (Brown et al., 2020) to investigate the patterns of certain and uncertain questions. Specifically, we query gpt-3.5-turbo five times with Chain-of-Thought prompts (Wei et al., 2023) with a temperature of 0.7, and calculate the entropy of the answers toward the same question (Diao et al., 2023b). If the model provides many different answers to the same question, the entropy should be high. Otherwise, the entropy should be low. The results are shown in Table 7. We observe that the average entropy of the answers on certain data D1D_{1} is lower than the entropy of uncertain data D0D_{0} data in most cases, which illustrates that when fed with certain questions, gpt-3.5-turbo is more likely to generate consistent answers. It will generate hallucinated answers to uncertain questions with much higher chances.

Therefore, we can conclude that R-Tuning divides the data into two folds. The uncertain questions are generally more difficult than certain questions because gpt-3.5-turbo’s answers vary more with the uncertain data. R-Tuning endows the model with abilities to identify and differentiate the difficulties of the questions. Therefore, our fine-tuned model becomes proactive in answering easy questions with certainty while being conservative in answering difficult questions, which eventually increases the precision and prevents the fine-tuned model from making too many mistakes.

Related Work

In this section, we review the progress on hallucinations of large language models (LLMs) and their uncertainty quantification methods.

Despite the outstanding performance of large language models with high fluency and coherence, they are still likely to hallucinate unfaithful and nonfactual facts (Maynez et al., 2020b). Recently, a variety of works have been done towards hallucination detection and mitigation. For hallucination detection, Lee et al. (2023) propose a benchmark for measuring the factuality of generation, using factual and nonfactual prompts. Manakul et al. (2023) propose SelfCheckGPT, a zero-resource, sample-based approach to detect hallucination by checking the consistency of multiple responses from LLM. For hallucination control, retrieval-augmented methods (Peng et al., 2023; Xie et al., 2023; Yue et al., 2023; Lyu et al., 2023) have shown effectiveness in mitigating the hallucination. Other methods, such as corruptions denoising (Chen et al., 2023), low-confidence validation (Varshney et al., 2023), multi-agent debate (Du et al., 2023), question-knowledge alignment (Zhang et al., 2023b), knowledge injection and teacher-student model (Elaraby et al., 2023), also improve the factuality of generation from multiple perspectives. Previous studies show the importance of the early discovery of hallucination, since the model will generate more incorrect statements to justify the answer if the previous answer is incorrect (Zhang et al., 2023a). In addition, Huang et al. (2023) found that LLMs cannot rectify themselves with their initial capabilities, displaying the importance of fine-tuning and external feedback.

Although large language models might have been taught to be honest during the pre-training, and instruction tuning on human preferred text (Ouyang et al., 2022; Bai et al., 2022), they still fail to recognize what and when they do not know. Our proposed method instructs the model to be aware of its knowledge gap between the instruction tuning datasets and the parametric knowledge, so that it possesses the refusal ability when it encounters instructions out of its knowledge.

2 Uncertainty Quantification of LLMs

Uncertainty quantification is a long-standing problem in machine learning. In the deep learning era, Guo et al. (2017) first identify the predictive confidence (a.k.a, predictive probability) of deep neural network lack of calibration in terms of the ECE metric (Expected Calibration Error) (Naeini et al., 2015). Chen et al. (2022) further study the investigate the calibration problem of pre-trained large language models and observe the same miscalibration problem on large language models. Active-Prompt (Diao et al., 2023b) introduces uncertainty to select questions for chain-of-thought annotation and demonstrates its effectiveness in actively and judiciously selecting and annotating the most helpful exemplars for in-context learning of LLMs.

Conclusion

In this paper, we propose a simple yet effective method, R-Tuning, to teach large language models to refuse unknown questions. It identifies the difference between the instruction tuning data and parametric knowledge of the model and splits the training data into a certain part and an uncertain part. Then, R-Tuning constructs the refusal-aware data by appending uncertainty expressions to the uncertain part. Empirically, R-Tuning outperforms the traditional instruction tuning baseline regarding AP score, illustrating a good trade-off between precision and recall. R-Tuning not only shows the refusal ability on in-domain data but also demonstrates such ability could be generalized to unseen tasks well. It displays that refusal is a fundamental ability and could be abstracted via multi-task learning, so we call it meta-skill. Further analysis of the perplexity and uncertainty of the training datasets reveals the rationale of our proposed method.

References

Appendix A Appendix

We conduct our experiments on nine datasets, which are described as follows.

ParaRel (Elazar et al., 2021): a dataset of factual knowledge with various prompts and relations that are originally for mask prediction. To align the dataset with the requirements of our auto-regressive models, we first change the format into question-answering and our models read the questions and generate the answers. Then, duplicated prompts of different templates but with the same entities are omitted for our question-answering task. It finally comes up with 25133 prompt-answer pairs of 31 domains. We split the ParaRel into two sets - the first 15 domains as in-domain data and the last 16 domains as out-of-domain data. We also equally split the in-domain data into training data and test data.

MMLU (Hendrycks et al., 2021): MMLU covers 57 tasks including mathematics, computer science, history, law, and more, which requires extensive world knowledge and problem-solving ability. The dataset is of multiple-choice format, and we can directly use it in our experiments.

WiCE (Kamoi et al., 2023): WiCE is a natural language inference (NLI) dataset for textual entailment. Each data sample consists of evidence and a claim, and the model should decide whether the evidence supports, partially supports, or doesn’t support the claim. We turn the dataset into multiple-choice questions with 3 choices for each question.

HotpotQA (Yang et al., 2018): HotpotQA is a question-answering dataset that requires complex reasoning among documents. We evaluate by providing the context documents and questions to see if the model can answer them. Since the test set of HotpotQA requires answer submission, we instead use the development set to do the evaluation.

FEVER (Thorne et al., 2018): FEVER is a dataset containing claims and supporting knowledge. The claims are classified as SUPPORTED, REFUTES, or NOT ENOUGH INFO. We turn it into a multiple-choice NLI task.

SelfAware (Yin et al., 2023): a dataset containing both answerable questions and unanswerable questions. We evaluate the unanswerable questions. It is expected to see our finetuned models refusing the unanswerable questions while other baselines do not possess such ability.

HaluEval (Li et al., 2023a): HaluEval is a dataset containing question-answering, dialogue, summarization, and user-query with correct answers and hallucinated answers. We only take the question-answering part.

FalseQA: FalseQA is a new open-domain dataset with questions inconsistent with common sense. There are no correct answers to the questions.

NEC: NEC is also a new open-domain dataset with questions containing some make-up concepts. There are also no correct answers to the questions.

Details about the original datasets are shown in Table 8. In Figure 6, we present the distribution of constructed refusal-aware data D0D_{0} and D1D_{1}.

A.2 Case Studies of R-Tuning-R

In this section, we display the detailed statistics in Table 9, and illustrate more case studies of R-Tuning-R in Table 10.

A.3 Uncertainty Text

In this section, we list the 16 uncertainty expressions from Yin et al. (2023):

There is no concrete answer to this question.

There is no public information available.

A.4 Min-Loss Training

Compared with the append verbalizer, replace verbalizer (e.g., R-Tuning-R) is a clear-cut way of producing uncertainty expressions by throwing away valuable labels which could potentially be leveraged for training. In address of this dimension of concern, we consider a modified cross entropy learning objective that pushes up the correct answer token and keeps the uncertainty expressions as the second most probable token choice. We call it min-loss training, which is optimized by gradient descent over the min loss between guessing the correct answer or just uncertainty expressions. It is formulated as follows:

where LL denotes the cross-entropy loss. To do so, we split the training data in half and adopt a two-stage training strategy. In the first stage, we train our model using the original method where the prompt template uses The answer is {ground-truth} if the model answers correctly, otherwise The answer is unknown. Once the model learns such a pattern after the first training stage, we calculate the min-loss with the equation 12. We only consider the loss of the unknown and the ground-truth label, and we mask the tokens before them. Since the ground-truth label may consider more than one token, we calculate the loss for the first token.

We evaluate the performance of min-loss strategy on the ParaRel dataset, and the results are shown in Table 11. It shows min-loss training outperforms R-Tuning-R in small models and in-domain settings. However, it underperforms R-Tuning-R in out-of-domain test sets. We also notice that in out-of-domain test sets, the accuracy of the model of 3B size is nearly the same as 7B’s and 13B’s. We identify such issues as a trade-off between the accuracy and the answer rate. When the model is proactive in answering more questions, it will inevitably make more mistakes. As the intrinsic parametric knowledge of the model is limited, there is no method to fine-tune a model with both high accuracy and a high answer rate.

A.5 Confidence Distribution of Training Dataset

We calculate the confidence of the certain data D1D_{1} and uncertain data d0d_{0}, and they are shown in Figure 7.

A.6 AP Scores of Each Dataset and Model Size with Figures

We calculate the AP scores for each dataset with different model sizes in multi-task experiments. The results are shown in Table 12 and Figure 8.