Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, Jonathan Herzig
Introduction
Pre-training Large Language Models (LLMs) on textual corpora embeds substantial factual knowledge in their parameters Petroni et al. (2019); AlKhamissi et al. (2022); Cohen et al. (2023), which is essential for excelling in various downstream applications. These models often require further alignment to desired behaviors, typically achieved through supervised fine-tuning on instruction-following tasks Wei et al. (2022); Mishra et al. (2022) and preference learning from human feedback Ouyang et al. (2022); Rafailov et al. (2024).
In the fine-tuning phase, the model is usually trained on outputs created by human annotators or other LLMs. As a result, the model may encounter new factual information, extending beyond the knowledge it acquired during pre-training. This raises the question of how LLMs integrate new facts outside of their pre-existing knowledge. One possibility is that the model simply adapts by learning this new factual information. However, a common conjecture posits that such exposure to new knowledge may encourage the model to hallucinate factually incorrect responses, as the model is essentially trained to generate facts that are not grounded in its pre-existing knowledge Schulman (2023); Huang et al. (2023); Gao (2021); Goldberg (2023); Gudibande et al. (2023).
In this work, we study how learning new factual knowledge through fine-tuning impacts the model’s tendency to hallucinate w.r.t. its pre-existing knowledge, exploring the above conjecture.While we focus on supervised fine-tuning, our findings are relevant to offline preference optimization methods such as DPO Rafailov et al. (2024) that may add new knowledge.
To study the impact of new knowledge, we must be able to assess whether a single fine-tuning example is consistent with the model’s knowledge. We propose SliCK, a hierarchy of four knowledge categories, derived from a continuous measure that quantifies the agreement between model-generated answers and the ground-truth labels. In SliCK, examples are first categorized into and types, where the latter corresponds to examples with facts that are most likely unknown to the model. The examples are subsequently split into three categories: , , and (Figure 2).
Equipped with the above method, we carefully design a controlled study, focused on closed-book question answering (QA), where we vary the proportion of the fine-tuning examples categorized as , while controlling for other factors.
Our study empirically demonstrates that learning from fine-tuning examples is linearly correlated with the model’s tendency to hallucinate w.r.t. its pre-existing knowledge (§4). Conversely, learning from examples is correlated with better utilization of pre-existing knowledge.
Through an analysis of the training dynamics, we discover that the LLM fits fine-tuning examples substantially slower than examples (top plot in Figure 1). This indicates that during fine-tuning, LLMs struggle to integrate new factual knowledge (present in the fine-tuning examples). Instead, they mostly learn to expose their pre-existing knowledge (using the fine-tuning examples).
From a practical perspective, mitigating overfitting using early-stopping (vertical dotted line in Figure 1) can minimize the risk of the hallucinations caused by fitting the examples, since they primarily emerge in later training stages as a form of overfitting (as illustrated by the development performance decline in the bottom plot of Figure 1). Alternatively, we also show that filtering-out the fine-tuning examples substantially reduces the risk of overfitting, without sacrificing performance.
We further evaluate the impact of fine-tuning examples from each of our three knowledge categories on performance (§5). Unexpectedly, we find that a model fine-tuned only on examples from the highest knowledge degree, denoted , does not yield the best results. Our analysis reveals that incorporating fine-tuning examples, representing facts with lower degrees of certainty, plays an important part in properly handling such examples in test time. This indicates that the composition of fine-tuning examples significantly influences the extent to which LLMs effectively utilize their pre-existing knowledge.
To summarize, we study the effect of new factual knowledge in the fine-tuning data by designing a controlled setup that isolates this factor. We find that fine-tuning examples that introduce new knowledge are learned slowly, which suggests that LLMs struggle to integrate new knowledge through fine-tuning and supports the view that LLMs mostly acquire knowledge through pre-training Zhou et al. (2023); Lin et al. (2023). However, we also find that as the model eventually learns new knowledge through fine-tuning, it becomes more prone to hallucinations w.r.t. its pre-existing knowledge. Collectively, our findings highlight the potential for unintended consequences when introducing new knowledge through fine-tuning, and imply that fine-tuning may be more useful as a mechanism to enhance the utilization of pre-existing knowledge.
Study Setup
Given a fine-tuning dataset and a pre-trained LLM , we denote by a model obtained by fine-tuning on . To study how new knowledge in affects ’s performance, we design a controlled setup creating variants of with varying proportions of examples that are unknown to .
When constructing , our objective is to reflect instruction tuning on diverse knowledge-intensive tasks while maintaining control over the experimental setting. We thus focus on factual knowledge that can be structured as (subject, relation, object) triplets, which are converted into closed-book QA format. In this setup, , where is a knowledge-seeking question corresponding to a specific triplet (e.g., “Where is Paris located?”) and is the ground-truth answer (e.g., “France”). To this end, we use EntityQuestions Sciavolino et al. (2021), where triplets from a diverse set of relations from Wikidata Vrandečić and Krötzsch (2014) are converted to QA pairs. These relations encompass a broad spectrum of factual knowledge, including biographical information, geographical data, ownership and authorship details, history and more. We use the original development and test splits, and we sub-sample the train split to create different variants of . We focus on 12 diverse relations and reserve 7 additional relations for an out-of-distribution test set, used (only) in §4.5.
As , we use the PaLM 2-M base model Anil et al. (2023). We focus on exact match (EM) as our evaluation metric.We validated that in our setting EM strongly correlates with word-level F1 Rajpurkar et al. (2016), and we choose EM as it is more intuitive for the purposes of our analysis. Full technical details are in §A.
Quantifying Knowledge in LLMs
To assess the effect of new knowledge in on the performance of , we have to annotate each pair in w.r.t. whether knows that the answer to is . To estimate this, we define a continuous measure based on samples from , and use it to divide pairs into four knowledge categories. We name this approach SliCK (Sampling-based Categorization of Knowledge).
We adopt the perspective that knows that the answer to is if it generates when prompted to answer Kadavath et al. (2022); Manakul et al. (2023). Since is a base model that has not been specifically fine-tuned to follow instructions, we prompt using in-context learning with few-shot exemplars. Following Rubin et al. (2022), we make sure that the few-shot exemplars have high semantic similarity to .In our study we achieve this by using exemplars from the same relation. E.g., if “Where is Paris located?”, the exemplars would follow the pattern “Where is {X} located?”.
In practice, can predict different answers since (1) the choice of exemplars influences individual predictions and (2) temperature sampling, if used, introduces randomness. To reflect this, we define as an estimate of how likely is to accurately generate the correct answer to , when prompted with random few-shot exemplars and using decoding temperature .
For the purposes of our study we approximate the value of using different random 4-shot prompts.We use 4-shot simply since we found it enough for to output answers in the correct format. For each 4-shot prompt, we predict the greedy answer using and sampled answers using . is estimated by the fraction of correct greedy answers, and by the fraction of correct sampled answers. Full details are in §C.
We define the category (bottom row in Figures 1(a) and 1(b)) to represent pairs for which never predicts the correct answer to . In our notations this means that . Alternatively, if , i.e. sometimes predicts the correct answer to , we consider as . In this choice, we posit that if prompting to answer can sometimes result with the correct answer , then must have some association with the relevant fact.
Recognizing that knowledge can vary in degrees of certainty and extent, we divide the pairs into three distinct categories (top three rows in Tables 1(a) and 1(b)). Motivated by the principle that should consistently predict if is , we put emphasis on greedy decoding outcomes, represented with . represents pairs for which always greedily predicts . If sometimes (but not always) greedily predicts , we consider as . Lastly, if never greedily predicts , we classify as .
We apply SliCK to annotate each pair in our dataset with its knowledge category w.r.t. .In EntityQuestions we have , , , , and . Full per-relation statistics are in §D. We analyze the quality of our categories in §6.
How Harmful are 𝚄𝚗𝚔𝚗𝚘𝚠𝚗𝚄𝚗𝚔𝚗𝚘𝚠𝚗\mathtt{Unknown} Examples?
In this section we study the effect of new knowledge in the fine-tuning dataset on performance. To isolate this effect, we vary the proportion of examples in , while controlling for other factors. Specifically, we fix and create variants of with of and examples (full details in §E). We treat the categories collectively (see Table 1(a)), and provide a per-category analysis in §5. We denote early-stopping based on the development set as early_stop (happens after 5-10 epochs) and 50 fine-tuning epochs as Convergence, as at this point always completely fits (i.e. training accuracy). We measure test performance as a proxy for hallucinations since we are in a closed-book QA setup with disjoint train/test splits, where the model has to use its per-existing knowledge to answer test questions (see §B for further discussion).
Figure 3(a) presents the performance as a function of the % of examples in , for different fine-tuning durations. Higher % leads to performance degradation, regardless of the fine-tuning duration, which indicates that examples are less useful than . Performance is also strongly affected by the fine-tuning duration, with early_stop typically yielding the best performance. Training for more epochs usually reduces performance (with the lowest performance observed for Convergence), which can be attributed to overfitting . Interestingly, this effect increases with larger (the inter-line spacing from early_stop exhibits a monotonic increase along the positive x-axis), suggesting that a higher % increases the risk of overfitting.
2 𝚄𝚗𝚔𝚗𝚘𝚠𝚗𝚄𝚗𝚔𝚗𝚘𝚠𝚗\mathtt{Unknown} Examples: Harmful or Neutral?
Since is fixed, performance drops for higher % could stem from simply the lower number of the fine-tuning examples. Thus, it is still not clear if examples are harmful or neutral. To address this, we measure the effect of filtering-out all the examples from . For each variant, we create a corresponding ablated variant , consisting only from the examples in . E.g., if has , we filter them out and are left with the remaining examples and get .
Figure 3(b) presents the results. Perhaps surprisingly, for early_stop the results for are almost identical to , indicating that the examples had a neutral effect on performance (as their removal had minimal impact). Conversely, the Convergence results show that with longer training, examples are actually very harmful. In this case under-performs , and the gap between them is proportional to the ratio.
Interestingly, for , the gap between early_stop and Convergence is very small (dotted lines), while this gap is very large for (full lines). This indicates that the presence of examples is what makes the variants with higher ratios more prone to overfitting.
We showed that examples are harmful, but their negative effect is mostly materialized in later training stages, and thus can be empirically avoided using early stopping. To better understand these trends, we analyze the training dynamics by examining which fine-tuning examples in were fitted by during various fine-tuning stages. Figure 1 presents the train accuracy of the and subsets of as a function of the fine-tuning duration. The development accuracy is presented in a zoomed-in plot at the bottom, as it falls within a narrower range. We include a breakdown of the train accuracy per category in §F.
fits fine-tuning examples substantially slower than . In early_stop (vertical dotted line), reaches peak performance on the development set, while fitting the majority of the examples but only a small fraction of the . In Figure 4, we show that this behavior is consistent across all our variants of . This can explain why in early_stop the examples had a neutral effect on performance (§4.2), as at this point still did not fit most of them. Lastly, since examples are the ones that are likely to introduce new factual knowledge, their significantly slow fitting rate suggests that LLMs struggle to acquire new factual knowledge through fine-tuning, instead they learn to expose their pre-existing knowledge using the examples.
Figure 1 demonstrates that after the development performance peaks at early_stop (vertical dotted line), it deteriorates as gradually fits more examples. In this section, we aim to characterize this relationship more accurately by assessing whether a simple linear dependency can tie the impact of fitting and training examples on test accuracy. To this end we use the following linear regression model:
where and are the number of the and examples in that fits.
We estimate the coefficientsFull details in §G. We note that this linear model is only valid in bounded region of , . by collecting (, , ) values after each epoch from models fine-tuned on all variants. Table 1 presents the results (top row). The high indicates a strong linear relationship between test accuracy and the type of training examples that are fitted. Our model entails that fitting examples hurts performance (), while fitting examples improves it (). The estimated negative impact from roughly matches the positive impact from ().
5 Generalization to New Relations
In the above setup, the pairs in the test set correspond to triplets with the same set of 12 relations appearing in . We now investigate whether our observed dynamics has a broader effect on the model’s knowledge, and transfers to relations not represented in . To test this, we reserve a subset of the relations for an out-of-distribution (OOD) test set, excluding them from the train and development splits. See §A for details and Tables 4 and 5 for in-distribution vs OOD relations.
Our results on the OOD test set reveal similar key insights: (1) Higher ratio leads to lower OOD test performance and (2) examples are harmful for OOD performance, but mostly when fits them. A linear model of the OOD test accuracy (Equation 1), shows similar trends: , , and (see Table 1). More details are in §H.
Overall, our insights transfer across relations. This essentially shows that fine-tuning on examples such as "Where is [E1] located?", can encourage hallucinations on seemingly unrelated questions, such as "Who founded [E2]?". This further supports the conclusion that the observed effects likely stem from the model learning the behavior of generating answers that are not grounded in its pre-existing knowledge.
Understanding Knowledge Types: Their Value and Impact
When addressing our main research question on the effect of fine-tuning examples, we treated the categories collectively for simplicity (see Table 1(a)). We now examine the effect of each category, exploring the following questions: Q1: How training examples from each category impact the test performance? Q2: What is the model’s performance across test examples from each category? To address Q1 we created single-category variants of the fine-tuning dataset . A variant of consisting solely of examples from the category is denoted as . For reference, we include a variant with the natural categories distribution in EntityQuestions, denoted . is fixed and identical to our experiments in §4. To address Q2, we further break down the test set performance by category. Table 2 presents the results.
Since examples are harmful, one might expect that it would be best to fine-tune on the most exemplary examples. Surprisingly, does not obtain the best overall results, as it excels on test examples, yet its performance on the remaining categories is inferior. yields the best overall performance. Compared to , enhances ’s performance on (), without compromising performance on (). This suggests that fine-tuning examples are essential for to correctly handle such examples during inference. It also demonstrates that with the right fine-tuning examples, becomes more capable of utilizing its pre-existing knowledge.
Limited Knowledge Enhances Overfitting.
In §4.2, we demonstrated that fine-tuning examples increase the risk of overfitting. We now observe that this also applies to , though to a lesser degree. Specifically, at Convergence, and experience significant performance drops compared to early_stop ( and ). With training to Convergence, they show a modest improvement on and but substantially degrade on and . This highlights that the decrease in performance is strongly attributed to an increased rate of hallucinations w.r.t. facts that were already known to after pre-training.
Interestingly, performs on-par with in early_stop, suggesting that the mere presence of examples in suffices for high performance on , even if has additional examples from other categories. Yet, ’s performance degrades significantlySee §I for details about this statistic significance test. after Convergence, under-performing – indicating that it still suffers from overfitting, most-likely due to the presence of and examples. Taken together these results demonstrate that stands out both in terms of top performance and reduced risk to overfitting.
SliCK Knowledge Categories Analysis
Assessing a model’s knowledge remains an open problem, particularly since evaluating the quality of such methods is challenging due to the lack of ground truth about what the model truly knows. In this work we proposed SliCK (§3): a four-category classification of facts w.r.t. the model’s knowledge. We now further analyze and discuss our design choices, hoping that SliCK can serve as a useful taxonomy to guide future research on this subject.
We first reflect on whether our choice of splitting into more fine-grained categories, based on the greedy decoding outcome, has been proven meaningful. As shown in Table 2, indeed captures facts with high degree of knowledge, as it consistently exceeds accuracy post fine-tuning, while and seem to represent weaker knowledge degrees. As intended, the performance on is worse that on but better than on . Additionally, the exact categories distinction we made was proven useful since it revealed important insights on the importance of the fine-tuning examples, as discussed in detail in §5.
Benchmarking Unknown Test Examples
A desired property for pairs classified as that appear in the test set, is that will incorrectly answer post fine-tuning (otherwise they are not truly ). Since in our closed-book QA setup the train and test sets are disjoint, the model has to rely on its pre-existing knowledge to answer test questions. In Table 2 we can see that the accuracy on is extremely low ( or less), which is a strong indicator that most of the examples are actually unknown to .
As a case study for comparison, we analyze the P(True) approach by Kadavath et al. (2022): a continuous score that estimates the probability a model assigns to the correctness of a specific answer. P(True) was originally used for self-evaluating model-generated answers, while we use it to assess whether considers the ground-truth answer as correct. In Figure 5, we explore classifying examples below a P(True) threshold as and compare this methodology to SliCK. Our results indicate that, at least in our setting, our approach categorizes examples for which the model’s performance after fine-tuning is significantly worse. Specifically, looking at fixed values on the x-axis shows that if we would label a similar fraction of test examples as using both methods, the accuracy on the P(True)-based examples would be much higher post fine-tuning.This is a preliminary analysis, and we leave a comprehensive comparison for future work. More details in §J. Lastly, the blue line shows that using samples from multiple few-shot prompts to approximate is crucial, as using leads to higher test accuracy on SliCK examples.
Discussion
This work highlights the risk in using supervised fine-tuning to update LLMs’ knowledge, as we present empirical evidence that acquiring new knowledge through fine-tuning is correlated with hallucinations w.r.t pre-existing knowledge. Additionally, this work raises important questions for future exploration regarding fine-tuning practices. We saw that examples are fitted slower than the ones, thus their negative effect manifests as a form of overfitting, which emphasizes the importance of using early-stopping instead of a fixed number of fine-tuning steps. However, early-stopping may be less effective when fine-tuning on numerous tasks with distinct optimal stopping points. An alternative solution can be to align the fine-tuning data with the model’s knowledge by filtering-out examples. We show initial evidence that this can reduce the risk of overfitting without compromising performance. A possible drawback of filtering is that fine-tuning examples can still be useful to teach LLMs to express uncertainty on test examples Zhang et al. (2023). This raises the question: can re-labeling fine-tuning examples with uncertainty expressions (e.g., “I don’t know”) reduce their negative effect? Our preliminary experiment (described in §K) suggests that the answer is yes, which indicates that such approaches could be the most promising. Exploring this is an interesting direction for future work.
Superficial Alignment Hypothesis.
Zhou et al. (2023) hypothesized that the knowledge and capabilities of LLMs are mostly learned during pre-training, while alignment is a simple process where the model learns the style or format for interacting with users. They substantiate this hypothesis by showing that fine-tuning on just high-quality examples can result with a competitive assistant LLM, named LIMA. As discussed in §4.3, we show evidence that LLMs struggle to acquire new knowledge present in the examples and mostly learn to utilize their pre-existing knowledge. We also showed that fine-tuning on examples led to sub-optimal utilization of pre-existing knowledge, despite our task format being simpler than LIMA’s and our dataset being six times larger. Taken together, our findings suggest that even though most of the LLM’s knowledge is indeed acquired through pre-training, the model learns more than just style or format through fine-tuning, as the selection of fine-tuning examples significantly influences the model’s capability to utilize its pre-existing knowledge post fine-tuning.
Related Work
Schulman (2023), Goldberg (2023) and Gudibande et al. (2023) mention the conjecture that fine-tuning on new factual knowledge may encourage hallucinations. Huang et al. (2023) categorized hallucination causes and formally defined this scenario as capability misalignment. They highlight that limited research addresses capability misalignment due to the challenge of defining the knowledge boundary of LLMs. Kang et al. (2024) showed that when a fine-tuned LLM encounters unknown queries at test time, its responses mimic the responses associated with the unknown examples in the fine-tuning data. Yin et al. (2023) showed that LLMs’ performance is not satisfactory when they face new knowledge in their input contexts and Lee et al. (2023) analyzed the impact of unknown in-context learning examples. To the best of our knowledge, our work is the first to empirically assess the impact of exposure to new knowledge through fine-tuning on tendency of the fine-tuned model to hallucinate.
Quantifying knowledge in LLMs.
SliCK can be seen as a confidence elicitation method for the ground truth label ( knows if it is confident that is the answer to ). Existing work derive calibrated confidence from LLMs by examining agreement across multiple samples Kuhn et al. (2023); Manakul et al. (2023); Tian et al. (2023a); Lyu et al. (2024), probing internal representations Azaria and Mitchell (2023); Burns et al. (2022), eliciting verbalized probability Tian et al. (2023b) or direct prompting Kadavath et al. (2022). Kadavath et al. also trained a separate P(IK) model to predict if the LLM knows the answer to . The label for P(IK) was approximated by the fraction of correct sampled answers, which is conceptually aligned with (§3). A key difference is that we also define the SliCK categories, and provide evidence that we capture meaningful and useful categories.
Conclusion
We study the impact of integrating new factual knowledge through fine-tuning on the model’s tendency to hallucinate. We first propose SliCK, a categorization of facts w.r.t. LLM’s knowledge. We then design a controlled study where we isolate the impact of new knowledge and rigorously evaluate its effects. We provide multiple insights on the fine-tuning dynamics, with the following key findings: (1) Acquiring new knowledge via supervised fine-tuning is correlated with hallucinations w.r.t. pre-existing knowledge. (2) LLMs struggle to integrate new knowledge through fine-tuning and mostly learn to use their pre-existing knowledge.
Limitations
Our experiments were conducted using a single LLM, and thus it is unclear whether results will vary with different LLMs. Having said that, our study is extremely compute-heavy and thus challenging to replicate on multiple LLMs: First, our fine-tuning is compute-heavy as its runs are very long as we wanted to analyze the behavior during different stages of fine-tuning (including the overfitting stages). Second, and most importantly, to facilitate our study we needed to annotate a large scale dataset w.r.t the SliCK categories. To derive reliable conclusions, it was crucial to accurately assess the model’s knowledge w.r.t. a single fine-tuning example. In our case we run 170 inference steps per example, i.e., more than inference steps to categorize our full dataset.
In addition, since we focus on closed-book QA, the practical implications from our study such as filtering-out fine-tuning examples still require validation in settings involving long-form text generation. To filter-out examples that introduce new factual knowledge in long-form generation tasks, one would need to make adaptations to SliCK and come up with an effective way to compare the sampled answer with the ground-truth to approximate . We leave this for future work. Long-form generation tasks introduce evaluation challenges, leading to a wide adoption of LLM-based evaluations. Our choice to focus explicitly on closed book QA facilitates more accurate evaluation that enhances the reliability of our findings.
Lastly, we did not test the effect of adding additional fine-tuning examples from diverse tasks into the fine-tuning mixture. While this could more closely approximate a typical instruction fine-tuning scenario, such dataset extension may introduce new factual knowledge in an uncontrollable way, which will limit our findings.
Acknowledgments
We would like to thank Ori Ram, Uri Shaham, Alon Jacovi, Mor Ventura, Yochai Blau, Eyal Ben-David, Avi Caciularu, Avinatan Hassidim and the members of Roi Reichart’s NLP group for reviewing the paper draft and providing valuable feedback. Special thanks to Uri Shaham for assisting in setting up the fine-tuning pipeline during the early stages of our research.
References
Appendix A Data Preprocessing
This section expands §2 with additional details about our data preprocessing steps. The EntityQuestions dataset Sciavolino et al. (2021) consists of train, development and test splits and spans 24 relations. Our train, development and test sets are curated based on the original splits from EntityQuestions. However, we use only 12 relations, since we wanted to reserve some relations for out-of-distribution test set. To avoid cherry-picking, the 12 relations used in our train, development and test sets are randomly sampled. The resulting relations are presented in Tables 3 and 4.
We reserved the remaining 12 relations for out-of-distribution test set. However, we found that in those 12 reserved relations, 5 were too similar to some of the relations that we train on (Table 3), thus we suspected that this could lead to a test set that is not truly out-of-distribution. To address that, we filtered out those relations and were left with 7 relations for our-of-distribution. Specifically we filtered-out the following relations:
P276 was filtered out since it directly overlaps with P131 since for both relations the question in EntityQuestions is of the form “Where is [E] located?”. P276 stands for “location” (https://www.wikidata.org/wiki/Property:P276) and P131 stands for “located in the administrative territorial entity” (https://www.wikidata.org/wiki/Property:P131).
P20, for which the question template is “Where did [E] die?”, was filtered out since it may require knowledge that relates to P19, for which the question template is “Where was [E] born?”. P20 stands for “place of death” (https://www.wikidata.org/wiki/Property:P20) and P19 stands for “place of birth” (https://www.wikidata.org/wiki/Property:P19).
P106, for which the question template is “What kind of work does [E] do?”, was filtered out since it may require knowledge that relates to P800, for which the question template is “What is [E] famous for?”. P106 stands for “occupation” (https://www.wikidata.org/wiki/Property:P106) and P800 stands for “notable work” (https://www.wikidata.org/wiki/Property:P800).
P413, for which the question template is “What position does [E] play?”, was filtered out since it may require knowledge that relates to P800, for which the question template is “What is [E] famous for?”. P413 stands for “position played on team / speciality” (https://www.wikidata.org/wiki/Property:P413) and P800 stands for “notable work” (https://www.wikidata.org/wiki/Property:P800).
P159, for which the question template is “Where is the headquarters of [E]?”, was filtered out since it may require knowledge that relates to P36, for which the question template is “What is the capital of [E]?”. P159 stands for “headquarters location” (https://www.wikidata.org/wiki/Property:P159) and P36 stands for “capital” (https://www.wikidata.org/wiki/Property:P36).
The 7 relations used for out-of-distribution test set are presented in Table 5.
Lastly, we perform two additional filtering steps: (1) To simplify the process of categorizing the examples w.r.t. ’s knowledge (§3), we filter-out examples with more than 1 correct answer. and of the EntityQuestions train and test set respectively. (2) We make sure that no subjects or objects overlap between the train and test sets, For example, the subject “Bruce Smith” appears with 2 different relations ( and ) yielding 2 examples: (“What kind of work does Bruce Smith do?”, “poet”) and (“Where was Bruce Smith born?”, “Faribault”). by filtering-out overlapping examples from the train set. of the EntityQuestions train set.
Appendix B Test performance as Proxy for Hallucinations
We now detail the relation between the test performance in our setting and hallucinations. In our study, poorer performance of a fine-tuned model , compared to another fine-tuned model on the test set, can be attributed to a higher rate of hallucinations in , relative to its pre-existing knowledge, due to the following explanation.
The test set can be conceptually divided into two types of questions. First, there are questions with answers that are unknown to . Those questions will remain unknown post fine-tuning, as we make sure that the training set is disjoint from the test set (§A). This means that both and will fail to answer these questions. Thus, the test performance difference between and is mostly attributed to the second type of questions: ones that are known to , i.e. can answer them correctly since it posses the relevant knowledge. Thus, and must rely on their pre-existing knowledge to answer such questions, and a lower performance on such question can be only categorized as an hallucination w.r.t. pre-existing knowledge.
This section expands §3 with additional details about our approximation. In our study we approximate based on the fraction of correct answers to sampled from . We begin with randomly sampling distinct -shot exemplars for each relation in our dataset (§A). Then, to approximate , we use to generate answers to using each of the exemplars from the relation corresponding to . We first use temperature sampling with to sample answers for each of the exemplars. is then approximated by the fraction of correct answers from the total of predictions. We also generate the greedy decoding prediction () for each of the exemplars. is then approximated by the fraction of correct answers from the total of predictions.Since we can only have one greedy prediction for every k-shot exemplars.
We use in our study, simply since we found it enough for to output answers in the correct format. We use and . The samples using are sampled from Top 40.
The exemplars are sampled from the development split. We sample different samples since we found that even when the few-shot exemplars are sampled per-relation, their exact choice still affects the prediction. In §6 and Figure 5 we show evidence that this also improves the quality of our categories.
Below is an example of our 4-shot prompt format, from real example from EntityQuestions with the relation representing occupation.https://www.wikidata.org/wiki/Property:P106 The question in this case is “What kind of work does Ron Konopka do?” and the ground truth asnwer is “geneticist”.
To decide whether a sampled answer is correct, we use the Exact Match (EM) metric to compare it with the ground truth answer. The main advantage in this choice is that when EM is True, we know that the answer is correct for . The main potential risk associated with this choice is that we may wrongly classify answers as incorrect due to paraphrases or answers with different granularity levels Wang et al. (2023); Kamalloo et al. (2023); Yona et al. (2024)). To address this, we perform an error analysis on 100 predictions for which EM was False. We randomly sample 50 greedy predictions () and 50 samples with . The results are in Table 6. This analysis suggest that in of the cases where EM is False, the predicted answer is indeed incorrect. Which is a reasonable performance for our purpose, especially considering that when EM is True the answer is correct.
Appendix D Data Annotation
we first calculate and for each pair in our preprocessed dataset (§2 and §A), using our approximation (§3 and §C). We then use these values to categorize each pair into one of our four categories (§3 and Figure 2). We provide the full statistics of the categories on the train and test set, as well as the out-of-distribution test set in Tables 3, 4 and 5.
Appendix E Fine-tuning Details
In §4 we examine the effect of new knowledge in the fine-tuning dataset on the performance of , by varying the proportion of examples in . When we create variants of with exactly of and examples, we make sure that the relation distribution remains consistent. To achieve that we sample of from each relation.
In §5 we create single-category variants of . Since we want to work with a fixed across all variants, we want to make sure that we have examples from each category. To ensure this, we measure the size of the smallest category in each relation (see the “Min” column in Table 3) and define as their sum. In other words, for each relation we calculate the size of the smallest category and sum these values. This leads to , as illustrated by the last column in Table 3. More formally, for each relation r in the training split, and for each category CAT from our 4 SliCK categories, we define to be the examples from category CAT and relation r. Consequently is the number of the examples in . For example (see Table 3). We then define:
where are the 12 relations from the training set.
Below is an example of our data format in the train, development and test sets, from real example from EntityQuestions with the relation representing occupation.https://www.wikidata.org/wiki/Property:P106 The question in this case is “What kind of work does Ron Konopka do?” and the ground truth asnwer is “geneticist”. Answer the following question. What kind of work does Ron Konopka do?
Fine-tuning hypeparameters.
We fine-tune every model for 50 epochs for all our model variants to completely fit the training set, so we can examine all stages of fine-tuning. We use learning rate of 1e-5, a batch size of 128, and a dropout rate of 0.05. We evaluate the models every epoch on the development set. The early_stop stopping criteria is defined to be the epoch with the maximum accuracy on the development set.
Appendix F Train Accuracy on Different 𝙺𝚗𝚘𝚠𝚗𝙺𝚗𝚘𝚠𝚗\mathtt{Known} Categories
In §4.3 we analyze the fine-tuning dynamic and present the training accuracy as function of the fine-tuning duration in Figure 1. For simplicity we treated the categories collectively. For reference we also include the plot with the full per-category breakdown in Figure 6.
Appendix G Linear Model
In §4.4 and §4.5 we use a linear model (Equation 1) that predicts that test accuracy and the out-of-distribution test accuracy. We estimate the parameters of this linear model based on results from all our variants of used in §4. For all these variants, we measure the test accuracy and the number of and fine-tuning examples that fits during different fine-tuning stages. This way we collect a dataset with examples of the form , which we use to fit a linear regression model.
Appendix H Out-of-distribution (OOD) Evaluation
In §4.5 we discuss out-of-distribution (OOD) results. In these experiments we simply used our OOD test set consisting of 7 relations unseen during fine-tuning (see §A). When we perform the analysis discussed in §4.1 and §4.2, we additionally evaluated the models on the OOD test set. For completeness, we add here Figure 7, which is the out-of-distribution version of Figure 3. Figure 7(a) presents the OOD test performance as a function of of examples in for different fine-tuning duration. The corresponding in-distribution results (Figure 3(a)) were discussed in §4.1. Figure 7(b) presents the OOD test performance for the ablation where we filter-out fine-tuning examples. The corresponding in-distribution results (Figure 3(b)) were discussed in §4.2. We notice that similar trends, just with a smaller overall magnitude of the performance drop, up to 6 points drop compared to up to 14 for in-distribution. This smaller drop magnitude is also reflected in smaller values of and (Table 1).
Appendix I Statistic Significance Tests
In §5 we present Table 2. As mentioned in the caption, we perform statistic significance tests for each column. To this end we compare all the values to the maximal value in this column.
For each subset of the test set, we randomly shuffle all the examples in it, split them up into 100 approximately equally sized subsets, and compute accuracy for each of them for all the models of interest. We then apply paired-sample t-test with and .
In Table 2, the best result is in bold, as well as all the results with statistically non-significant difference from the best with . We additionally include a copy of Table 2 where all the statistical tests outcomes are annotated, see Table 7. We can see that in almost all cases the difference is statistically significant with , except two cases where it is only with ( and ).
Since we also discuss “horizontal” comparisons, where we compare early_stop to Convergence, we additionally run significance tests (not annotated in Table 2) for , comparing early_stop to Convergence. The difference for was not statistically significant while for all others (including ) it was significant with .
Appendix J The P(True) Case Study
In §6 we used the P(True) metric from Kadavath et al. (2022) as a case study for comparison. In Figure 5 we compare our category vs classifying as based on a threshold of P(True). We calculated P(True) for every pair in the test set using Kadavath et al. (2022)’s prompt:
We then treated pairs with P(True) below a threshold as . We experimented with each possible threshold in $T\mathtt{Unknown}\mathtt{Unknown}N_{\text{ex}}=10N_{\text{ex}}D_{\mathtt{Natural}}$ (§5).
In this work we showed that fitting fine-tuning examples negatively affects the test performance. However, this negative effect manifests as a form of overfitting. From practical perspective, we showed that we can mitigate overfitting by either using early-stopping or filtering-out examples from the fine-tuning dataset.
We now perform a preliminary experiment where check whether fine-tuning the model to abstain from examples can also be a potential mitigation. Specifically, we replace the label of the fine-tuning examples with the expression “I don’t know” and test whether this mitigates the observed overfitting.
Table 8 presents the of the test questions that were answered (i.e. did not respond with “I don’t know”) and the accuracy on those questions. This experiment was conducted on the variant with . The first row is for the original result with as a reference and the second row is for the results with , where the ground-truth label of the of the examples in was replaced with “I don’t know”
Consistent with the findings from previous work Zhang et al. (2023), we observe an improved accuracy on willingly answered test examples (when comparing vs ). When we compare early_stop vs Convergence for we observe a performance drop () which illustrates the overfitting effect. However, we observe that re-labeling the examples with uncertainty expression seem to reduce the risk of overfitting. Specifically, the accuracy for remains for both early_stop and Convergence, with a small decrease on the number of willingly answered questions ()