Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?

Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, Jonathan Herzig

Introduction

Pre-training Large Language Models (LLMs) on textual corpora embeds substantial factual knowledge in their parameters Petroni et al. (2019); AlKhamissi et al. (2022); Cohen et al. (2023), which is essential for excelling in various downstream applications. These models often require further alignment to desired behaviors, typically achieved through supervised fine-tuning on instruction-following tasks Wei et al. (2022); Mishra et al. (2022) and preference learning from human feedback Ouyang et al. (2022); Rafailov et al. (2024).

In the fine-tuning phase, the model is usually trained on outputs created by human annotators or other LLMs. As a result, the model may encounter new factual information, extending beyond the knowledge it acquired during pre-training. This raises the question of how LLMs integrate new facts outside of their pre-existing knowledge. One possibility is that the model simply adapts by learning this new factual information. However, a common conjecture posits that such exposure to new knowledge may encourage the model to hallucinate factually incorrect responses, as the model is essentially trained to generate facts that are not grounded in its pre-existing knowledge Schulman (2023); Huang et al. (2023); Gao (2021); Goldberg (2023); Gudibande et al. (2023).

In this work, we study how learning new factual knowledge through fine-tuning impacts the model’s tendency to hallucinate w.r.t. its pre-existing knowledge, exploring the above conjecture.While we focus on supervised fine-tuning, our findings are relevant to offline preference optimization methods such as DPO Rafailov et al. (2024) that may add new knowledge.

To study the impact of new knowledge, we must be able to assess whether a single fine-tuning example is consistent with the model’s knowledge. We propose SliCK, a hierarchy of four knowledge categories, derived from a continuous measure that quantifies the agreement between model-generated answers and the ground-truth labels. In SliCK, examples are first categorized into Known\mathtt{Known} and Unknown\mathtt{Unknown} types, where the latter corresponds to examples with facts that are most likely unknown to the model. The Known\mathtt{Known} examples are subsequently split into three categories: HighlyKnown\mathtt{HighlyKnown}, MaybeKnown\mathtt{MaybeKnown}, and WeaklyKnown\mathtt{WeaklyKnown} (Figure 2).

Equipped with the above method, we carefully design a controlled study, focused on closed-book question answering (QA), where we vary the proportion of the fine-tuning examples categorized as Unknown\mathtt{Unknown}, while controlling for other factors.

Our study empirically demonstrates that learning from Unknown\mathtt{Unknown} fine-tuning examples is linearly correlated with the model’s tendency to hallucinate w.r.t. its pre-existing knowledge (§4). Conversely, learning from Known\mathtt{Known} examples is correlated with better utilization of pre-existing knowledge.

Through an analysis of the training dynamics, we discover that the LLM fits Unknown\mathtt{Unknown} fine-tuning examples substantially slower than Known\mathtt{Known} examples (top plot in Figure 1). This indicates that during fine-tuning, LLMs struggle to integrate new factual knowledge (present in the Unknown\mathtt{Unknown} fine-tuning examples). Instead, they mostly learn to expose their pre-existing knowledge (using the Known\mathtt{Known} fine-tuning examples).

From a practical perspective, mitigating overfitting using early-stopping (vertical dotted line in Figure 1) can minimize the risk of the hallucinations caused by fitting the Unknown\mathtt{Unknown} examples, since they primarily emerge in later training stages as a form of overfitting (as illustrated by the development performance decline in the bottom plot of Figure 1). Alternatively, we also show that filtering-out the Unknown\mathtt{Unknown} fine-tuning examples substantially reduces the risk of overfitting, without sacrificing performance.

We further evaluate the impact of fine-tuning examples from each of our three Known\mathtt{Known} knowledge categories on performance (§5). Unexpectedly, we find that a model fine-tuned only on examples from the highest knowledge degree, denoted HighlyKnown\mathtt{HighlyKnown}, does not yield the best results. Our analysis reveals that incorporating MaybeKnown\mathtt{MaybeKnown} fine-tuning examples, representing facts with lower degrees of certainty, plays an important part in properly handling such examples in test time. This indicates that the composition of fine-tuning examples significantly influences the extent to which LLMs effectively utilize their pre-existing knowledge.

To summarize, we study the effect of new factual knowledge in the fine-tuning data by designing a controlled setup that isolates this factor. We find that fine-tuning examples that introduce new knowledge are learned slowly, which suggests that LLMs struggle to integrate new knowledge through fine-tuning and supports the view that LLMs mostly acquire knowledge through pre-training Zhou et al. (2023); Lin et al. (2023). However, we also find that as the model eventually learns new knowledge through fine-tuning, it becomes more prone to hallucinations w.r.t. its pre-existing knowledge. Collectively, our findings highlight the potential for unintended consequences when introducing new knowledge through fine-tuning, and imply that fine-tuning may be more useful as a mechanism to enhance the utilization of pre-existing knowledge.

Study Setup

Given a fine-tuning dataset DD and a pre-trained LLM MM, we denote by MDM_{D} a model obtained by fine-tuning MM on DD. To study how new knowledge in DD affects MDM_{D}’s performance, we design a controlled setup creating variants of DD with varying proportions of examples that are unknown to MM.

When constructing DD, our objective is to reflect instruction tuning on diverse knowledge-intensive tasks while maintaining control over the experimental setting. We thus focus on factual knowledge that can be structured as (subject, relation, object) triplets, which are converted into closed-book QA format. In this setup, D={(qi,ai)}i=1ND=\{(q_{i},a_{i})\}_{i=1}^{N}, where qq is a knowledge-seeking question corresponding to a specific triplet (e.g., “Where is Paris located?”) and aa is the ground-truth answer (e.g., “France”). To this end, we use EntityQuestions Sciavolino et al. (2021), where triplets from a diverse set of relations from Wikidata Vrandečić and Krötzsch (2014) are converted to QA pairs. These relations encompass a broad spectrum of factual knowledge, including biographical information, geographical data, ownership and authorship details, history and more. We use the original development and test splits, and we sub-sample the train split to create different variants of DD. We focus on 12 diverse relations and reserve 7 additional relations for an out-of-distribution test set, used (only) in §4.5.

As MM, we use the PaLM 2-M base model Anil et al. (2023). We focus on exact match (EM) as our evaluation metric.We validated that in our setting EM strongly correlates with word-level F1 Rajpurkar et al. (2016), and we choose EM as it is more intuitive for the purposes of our analysis. Full technical details are in §A.

Quantifying Knowledge in LLMs

To assess the effect of new knowledge in DD on the performance of MDM_{D}, we have to annotate each (q,a)(q,a) pair in DD w.r.t. whether MM knows that the answer to qq is aa. To estimate this, we define a continuous PCorrectP_{\mathtt{Correct}} measure based on samples from MM, and use it to divide (q,a)(q,a) pairs into four knowledge categories. We name this approach SliCK (Sampling-based Categorization of Knowledge).

We adopt the perspective that MM knows that the answer to qq is aa if it generates aa when prompted to answer qq Kadavath et al. (2022); Manakul et al. (2023). Since MM is a base model that has not been specifically fine-tuned to follow instructions, we prompt MM using in-context learning with few-shot exemplars. Following Rubin et al. (2022), we make sure that the few-shot exemplars have high semantic similarity to qq.In our study we achieve this by using exemplars from the same relation. E.g., if q=q=“Where is Paris located?”, the exemplars would follow the pattern “Where is {X} located?”.

In practice, MM can predict different answers since (1) the choice of exemplars influences individual predictions and (2) temperature sampling, if used, introduces randomness. To reflect this, we define PCorrect(q,a;M,T)P_{\mathtt{Correct}}(q,a;M,T) as an estimate of how likely is MM to accurately generate the correct answer aa to qq, when prompted with random few-shot exemplars and using decoding temperature TT.

For the purposes of our study we approximate the value of PCorrectP_{\mathtt{Correct}} using Nex=10N_{\text{ex}}=10 different random 4-shot prompts.We use 4-shot simply since we found it enough for MM to output answers in the correct format. For each 4-shot prompt, we predict the greedy answer using T=0T=0 and 1616 sampled answers using T=0.5T=0.5. PCorrect(q,a;M,T=0)P_{\mathtt{Correct}}(q,a;M,T=0) is estimated by the fraction of correct greedy answers, and PCorrect(q,a;M,T>0)P_{\mathtt{Correct}}(q,a;M,T>0) by the fraction of correct sampled answers. Full details are in §C.

We define the Unknown\mathtt{Unknown} category (bottom row in Figures 1(a) and 1(b)) to represent (q,a)(q,a) pairs for which MM never predicts the correct answer to qq. In our notations this means that PCorrect(q,a;M,T≥0)=0P_{\mathtt{Correct}}(q,a;M,T\geq 0)=0. Alternatively, if PCorrect(q,a;M,T≥0)>0P_{\mathtt{Correct}}(q,a;M,T\geq 0)>0, i.e. MM sometimes predicts the correct answer to qq, we consider (q,a)(q,a) as Known\mathtt{Known}. In this choice, we posit that if prompting MM to answer qq can sometimes result with the correct answer aa, then MM must have some association with the relevant fact.

Recognizing that knowledge can vary in degrees of certainty and extent, we divide the Known\mathtt{Known} (q,a)(q,a) pairs into three distinct categories (top three rows in Tables 1(a) and 1(b)). Motivated by the principle that MM should consistently predict aa if (q,a)(q,a) is Known\mathtt{Known}, we put emphasis on greedy decoding outcomes, represented with PCorrect(q,a;M,T=0)P_{\mathtt{Correct}}(q,a;M,T=0). HighlyKnown\mathtt{HighlyKnown} represents (q,a)(q,a) pairs for which MM always greedily predicts aa. If MM sometimes (but not always) greedily predicts aa, we consider (q,a)(q,a) as MaybeKnown\mathtt{MaybeKnown}. Lastly, if MM never greedily predicts aa, we classify (q,a)(q,a) as WeaklyKnown\mathtt{WeaklyKnown}.

We apply SliCK to annotate each (q,a)(q,a) pair in our dataset with its knowledge category w.r.t. MM.In EntityQuestions we have 24%24\% HighlyKnown\mathtt{HighlyKnown}, 23%23\% MaybeKnown\mathtt{MaybeKnown}, 17%17\%, WeaklyKnown\mathtt{WeaklyKnown}, and 36%36\% Unknown\mathtt{Unknown}. Full per-relation statistics are in §D. We analyze the quality of our categories in §6.

How Harmful are 𝚄𝚗𝚔𝚗𝚘𝚠𝚗𝚄𝚗𝚔𝚗𝚘𝚠𝚗\mathtt{Unknown} Examples?

In this section we study the effect of new knowledge in the fine-tuning dataset DD on performance. To isolate this effect, we vary the proportion of Unknown\mathtt{Unknown} examples in DD, while controlling for other factors. Specifically, we fix ∣D∣|D| and create variants of DD with X%X\% of Unknown\mathtt{Unknown} and (100−X)%(100-X)\% Known\mathtt{Known} examples (full details in §E). We treat the Known\mathtt{Known} categories collectively (see Table 1(a)), and provide a per-category analysis in §5. We denote early-stopping based on the development set as early_stop (happens after 5-10 epochs) and 50 fine-tuning epochs as Convergence, as at this point MM always completely fits DD (i.e. 100%100\% training accuracy). We measure test performance as a proxy for hallucinations since we are in a closed-book QA setup with disjoint train/test splits, where the model has to use its per-existing knowledge to answer test questions (see §B for further discussion).

Figure 3(a) presents the performance as a function of the % of Unknown\mathtt{Unknown} examples in DD, for different fine-tuning durations. Higher %Unknown\mathtt{Unknown} leads to performance degradation, regardless of the fine-tuning duration, which indicates that Unknown\mathtt{Unknown} examples are less useful than Known\mathtt{Known}. Performance is also strongly affected by the fine-tuning duration, with early_stop typically yielding the best performance. Training for more epochs usually reduces performance (with the lowest performance observed for Convergence), which can be attributed to overfitting DD. Interestingly, this effect increases with larger %\%Unknown\mathtt{Unknown} (the inter-line spacing from early_stop exhibits a monotonic increase along the positive x-axis), suggesting that a higher %Unknown\mathtt{Unknown} increases the risk of overfitting.

2 𝚄𝚗𝚔𝚗𝚘𝚠𝚗𝚄𝚗𝚔𝚗𝚘𝚠𝚗\mathtt{Unknown} Examples: Harmful or Neutral?

Since ∣D∣|D| is fixed, performance drops for higher %Unknown\mathtt{Unknown} could stem from simply the lower number of the Known\mathtt{Known} fine-tuning examples. Thus, it is still not clear if Unknown\mathtt{Unknown} examples are harmful or neutral. To address this, we measure the effect of filtering-out all the Unknown\mathtt{Unknown} examples from DD. For each DD variant, we create a corresponding ablated variant DKnownD_{\mathtt{Known}}, consisting only from the Known\mathtt{Known} examples in DD. E.g., if DD has 25%25\% Unknown\mathtt{Unknown}, we filter them out and are left with the remaining 75%75\% Known\mathtt{Known} examples and get ∣|DKnownD_{\mathtt{Known}} ∣=0.75×∣D∣|=0.75\times|D|.

Figure 3(b) presents the results. Perhaps surprisingly, for early_stop the results for DD are almost identical to DKnownD_{\mathtt{Known}}, indicating that the Unknown\mathtt{Unknown} examples had a neutral effect on performance (as their removal had minimal impact). Conversely, the Convergence results show that with longer training, Unknown\mathtt{Unknown} examples are actually very harmful. In this case DD under-performs DKnownD_{\mathtt{Known}}, and the gap between them is proportional to the Unknown\mathtt{Unknown} ratio.

Interestingly, for DKnownD_{\mathtt{Known}}, the gap between early_stop and Convergence is very small (dotted lines), while this gap is very large for DD (full lines). This indicates that the presence of Unknown\mathtt{Unknown} examples is what makes the variants with higher Unknown\mathtt{Unknown} ratios more prone to overfitting.

We showed that Unknown\mathtt{Unknown} examples are harmful, but their negative effect is mostly materialized in later training stages, and thus can be empirically avoided using early stopping. To better understand these trends, we analyze the training dynamics by examining which fine-tuning examples in DD were fitted by MM during various fine-tuning stages. Figure 1 presents the train accuracy of the Known\mathtt{Known} and Unknown\mathtt{Unknown} subsets of DD as a function of the fine-tuning duration. The development accuracy is presented in a zoomed-in plot at the bottom, as it falls within a narrower range. We include a breakdown of the train accuracy per Known\mathtt{Known} category in §F.

MM fits Unknown\mathtt{Unknown} fine-tuning examples substantially slower than Known\mathtt{Known}. In early_stop (vertical dotted line), MM reaches peak performance on the development set, while fitting the majority of the Known\mathtt{Known} examples but only a small fraction of the Unknown\mathtt{Unknown}. In Figure 4, we show that this behavior is consistent across all our variants of DD. This can explain why in early_stop the Unknown\mathtt{Unknown} examples had a neutral effect on performance (§4.2), as at this point MM still did not fit most of them. Lastly, since Unknown\mathtt{Unknown} examples are the ones that are likely to introduce new factual knowledge, their significantly slow fitting rate suggests that LLMs struggle to acquire new factual knowledge through fine-tuning, instead they learn to expose their pre-existing knowledge using the Known\mathtt{Known} examples.

Figure 1 demonstrates that after the development performance peaks at early_stop (vertical dotted line), it deteriorates as MM gradually fits more Unknown\mathtt{Unknown} examples. In this section, we aim to characterize this relationship more accurately by assessing whether a simple linear dependency can tie the impact of fitting Known\mathtt{Known} and Unknown\mathtt{Unknown} training examples on test accuracy. To this end we use the following linear regression model:

where NKnN_{\text{Kn}} and NUnkN_{\text{Unk}} are the number of the Known\mathtt{Known} and Unknown\mathtt{Unknown} examples in DD that MM fits.

We estimate the coefficientsFull details in §G. We note that this linear model is only valid in bounded region of Nkn≤∣D∣N_{\text{kn}}\leq|D|, Nunk≤∣D∣N_{\text{unk}}\leq|D|. by collecting (AccuracyAccuracy, NKnN_{\text{Kn}}, NUnkN_{\text{Unk}}) values after each epoch from models fine-tuned on all DD variants. Table 1 presents the results (top row). The high R2R^{2} indicates a strong linear relationship between test accuracy and the type of training examples that are fitted. Our model entails that fitting Unknown\mathtt{Unknown} examples hurts performance (βunk<0\beta_{unk}<0), while fitting Known\mathtt{Known} examples improves it (βkn>0\beta_{\text{kn}}>0). The estimated negative impact from Unknown\mathtt{Unknown} roughly matches the positive impact from Known\mathtt{Known} (∣βukn∣≈∣βkn∣|\beta_{\text{ukn}}|\approx|\beta_{\text{kn}}|).

5 Generalization to New Relations

In the above setup, the (q,a)(q,a) pairs in the test set correspond to triplets with the same set of 12 relations appearing in DD. We now investigate whether our observed dynamics has a broader effect on the model’s knowledge, and transfers to relations not represented in DD. To test this, we reserve a subset of the relations for an out-of-distribution (OOD) test set, excluding them from the train and development splits. See §A for details and Tables 4 and 5 for in-distribution vs OOD relations.

Our results on the OOD test set reveal similar key insights: (1) Higher Unknown\mathtt{Unknown} ratio leads to lower OOD test performance and (2) Unknown\mathtt{Unknown} examples are harmful for OOD performance, but mostly when MM fits them. A linear model of the OOD test accuracy (Equation 1), shows similar trends: βunk<0\beta_{\text{unk}}<0, βkn>0\beta_{\text{kn}}>0, ∣βukn∣≈∣βkn∣|\beta_{\text{ukn}}|\approx|\beta_{\text{kn}}| and R2=0.95R^{2}=0.95 (see Table 1). More details are in §H.

Overall, our insights transfer across relations. This essentially shows that fine-tuning on Unknown\mathtt{Unknown} examples such as "Where is [E1] located?", can encourage hallucinations on seemingly unrelated questions, such as "Who founded [E2]?". This further supports the conclusion that the observed effects likely stem from the model learning the behavior of generating answers that are not grounded in its pre-existing knowledge.

Understanding Knowledge Types: Their Value and Impact

When addressing our main research question on the effect of Unknown\mathtt{Unknown} fine-tuning examples, we treated the Known\mathtt{Known} categories collectively for simplicity (see Table 1(a)). We now examine the effect of each category, exploring the following questions: Q1: How training examples from each category impact the test performance? Q2: What is the model’s performance across test examples from each category? To address Q1 we created single-category variants of the fine-tuning dataset DD. A variant of DD consisting solely of examples from the category CAT\mathtt{CAT} is denoted as DCATD_{\mathtt{CAT}}. For reference, we include a variant with the natural categories distribution in EntityQuestions, denoted DNaturalD_{\mathtt{Natural}}. ∣D∣|D| is fixed and identical to our experiments in §4. To address Q2, we further break down the test set performance by category. Table 2 presents the results.

Since Unknown\mathtt{Unknown} examples are harmful, one might expect that it would be best to fine-tune on the most exemplary HighlyKnown\mathtt{HighlyKnown} examples. Surprisingly, DHighlyKnownD_{\mathtt{HighlyKnown}} does not obtain the best overall results, as it excels on HighlyKnown\mathtt{HighlyKnown} test examples, yet its performance on the remaining categories is inferior. DMaybeKnownD_{\mathtt{MaybeKnown}} yields the best overall performance. Compared to DHighlyKnownD_{\mathtt{HighlyKnown}}, DMaybeKnownD_{\mathtt{MaybeKnown}} enhances MDM_{D}’s performance on MaybeKnown\mathtt{MaybeKnown} (60.1 ⁣→ ⁣69.960.1{\mkern-3.0mu}\rightarrow{\mkern-3.0mu}69.9), without compromising performance on HighlyKnown\mathtt{HighlyKnown} (98.7→98.498.7{\mkern-2.0mu}\rightarrow{\mkern-2.0mu}98.4). This suggests that MaybeKnown\mathtt{MaybeKnown} fine-tuning examples are essential for MDM_{D} to correctly handle such examples during inference. It also demonstrates that with the right fine-tuning examples, MDM_{D} becomes more capable of utilizing its pre-existing knowledge.

Limited Knowledge Enhances Overfitting.

In §4.2, we demonstrated that Unknown\mathtt{Unknown} fine-tuning examples increase the risk of overfitting. We now observe that this also applies to WeaklyKnown\mathtt{WeaklyKnown}, though to a lesser degree. Specifically, at Convergence, DWeaklyKnownD_{\mathtt{WeaklyKnown}} and DUnknownD_{\mathtt{Unknown}} experience significant performance drops compared to early_stop (39.2 ⁣→ ⁣35.439.2{\mkern-3.0mu}\rightarrow{\mkern-3.0mu}35.4 and 37.5 ⁣→ ⁣25.837.5{\mkern-3.0mu}\rightarrow{\mkern-3.0mu}25.8). With training to Convergence, they show a modest improvement on WeaklyKnown\mathtt{WeaklyKnown} and Unknown\mathtt{Unknown} but substantially degrade on HighlyKnown\mathtt{HighlyKnown} and MaybeKnown\mathtt{MaybeKnown}. This highlights that the decrease in performance is strongly attributed to an increased rate of hallucinations w.r.t. facts that were already known to MM after pre-training.

Interestingly, DNaturalD_{\mathtt{Natural}} performs on-par with DMaybeKnownD_{\mathtt{MaybeKnown}} in early_stop, suggesting that the mere presence of MaybeKnown\mathtt{MaybeKnown} examples in DD suffices for high performance on MaybeKnown\mathtt{MaybeKnown}, even if DD has additional examples from other categories. Yet, DNaturalD_{\mathtt{Natural}}’s performance degrades significantlySee §I for details about this statistic significance test. after Convergence, under-performing DMaybeKnownD_{\mathtt{MaybeKnown}}– indicating that it still suffers from overfitting, most-likely due to the presence of WeaklyKnown\mathtt{WeaklyKnown} and Unknown\mathtt{Unknown} examples. Taken together these results demonstrate that DMaybeKnownD_{\mathtt{MaybeKnown}} stands out both in terms of top performance and reduced risk to overfitting.

SliCK Knowledge Categories Analysis

Assessing a model’s knowledge remains an open problem, particularly since evaluating the quality of such methods is challenging due to the lack of ground truth about what the model truly knows. In this work we proposed SliCK (§3): a four-category classification of facts w.r.t. the model’s knowledge. We now further analyze and discuss our design choices, hoping that SliCK can serve as a useful taxonomy to guide future research on this subject.

We first reflect on whether our choice of splitting Known\mathtt{Known} into more fine-grained categories, based on the greedy decoding outcome, has been proven meaningful. As shown in Table 2, HighlyKnown\mathtt{HighlyKnown} indeed captures facts with high degree of knowledge, as it consistently exceeds 95%95\% accuracy post fine-tuning, while MaybeKnown\mathtt{MaybeKnown} and WeaklyKnown\mathtt{WeaklyKnown} seem to represent weaker knowledge degrees. As intended, the performance on WeaklyKnown\mathtt{WeaklyKnown} is worse that on MaybeKnown\mathtt{MaybeKnown} but better than on Unknown\mathtt{Unknown}. Additionally, the exact categories distinction we made was proven useful since it revealed important insights on the importance of the MaybeKnown\mathtt{MaybeKnown} fine-tuning examples, as discussed in detail in §5.

Benchmarking Unknown Test Examples

A desired property for (q,a)(q,a) pairs classified as Unknown\mathtt{Unknown} that appear in the test set, is that MM will incorrectly answer qq post fine-tuning (otherwise they are not truly Unknown\mathtt{Unknown}). Since in our closed-book QA setup the train and test sets are disjoint, the model has to rely on its pre-existing knowledge to answer test questions. In Table 2 we can see that the accuracy on Unknown\mathtt{Unknown} is extremely low (3.2%3.2\% or less), which is a strong indicator that most of the Unknown\mathtt{Unknown} examples are actually unknown to MM.

As a case study for comparison, we analyze the P(True) approach by Kadavath et al. (2022): a continuous score that estimates the probability a model assigns to the correctness of a specific answer. P(True) was originally used for self-evaluating model-generated answers, while we use it to assess whether MM considers the ground-truth answer as correct. In Figure 5, we explore classifying examples below a P(True) threshold as Unknown\mathtt{Unknown} and compare this methodology to SliCK. Our results indicate that, at least in our setting, our approach categorizes Unknown\mathtt{Unknown} examples for which the model’s performance after fine-tuning is significantly worse. Specifically, looking at fixed values on the x-axis shows that if we would label a similar fraction of test examples as Unknown\mathtt{Unknown} using both methods, the accuracy on the P(True)-based Unknown\mathtt{Unknown} examples would be much higher post fine-tuning.This is a preliminary analysis, and we leave a comprehensive comparison for future work. More details in §J. Lastly, the blue line shows that using samples from multiple few-shot prompts to approximate PCorrectP_{\mathtt{Correct}} is crucial, as using Nex<10N_{\text{ex}}<10 leads to higher test accuracy on SliCK Unknown\mathtt{Unknown} examples.

Discussion

This work highlights the risk in using supervised fine-tuning to update LLMs’ knowledge, as we present empirical evidence that acquiring new knowledge through fine-tuning is correlated with hallucinations w.r.t pre-existing knowledge. Additionally, this work raises important questions for future exploration regarding fine-tuning practices. We saw that Unknown\mathtt{Unknown} examples are fitted slower than the Known\mathtt{Known} ones, thus their negative effect manifests as a form of overfitting, which emphasizes the importance of using early-stopping instead of a fixed number of fine-tuning steps. However, early-stopping may be less effective when fine-tuning on numerous tasks with distinct optimal stopping points. An alternative solution can be to align the fine-tuning data with the model’s knowledge by filtering-out Unknown\mathtt{Unknown} examples. We show initial evidence that this can reduce the risk of overfitting without compromising performance. A possible drawback of filtering is that Unknown\mathtt{Unknown} fine-tuning examples can still be useful to teach LLMs to express uncertainty on Unknown\mathtt{Unknown} test examples Zhang et al. (2023). This raises the question: can re-labeling Unknown\mathtt{Unknown} fine-tuning examples with uncertainty expressions (e.g., “I don’t know”) reduce their negative effect? Our preliminary experiment (described in §K) suggests that the answer is yes, which indicates that such approaches could be the most promising. Exploring this is an interesting direction for future work.

Superficial Alignment Hypothesis.

Zhou et al. (2023) hypothesized that the knowledge and capabilities of LLMs are mostly learned during pre-training, while alignment is a simple process where the model learns the style or format for interacting with users. They substantiate this hypothesis by showing that fine-tuning on just 1k\mathtt{1k} high-quality examples can result with a competitive assistant LLM, named LIMA. As discussed in §4.3, we show evidence that LLMs struggle to acquire new knowledge present in the Unknown\mathtt{Unknown} examples and mostly learn to utilize their pre-existing knowledge. We also showed that fine-tuning on HighlyKnown\mathtt{HighlyKnown} examples led to sub-optimal utilization of pre-existing knowledge, despite our task format being simpler than LIMA’s and our dataset being six times larger. Taken together, our findings suggest that even though most of the LLM’s knowledge is indeed acquired through pre-training, the model learns more than just style or format through fine-tuning, as the selection of fine-tuning examples significantly influences the model’s capability to utilize its pre-existing knowledge post fine-tuning.

Related Work

Schulman (2023), Goldberg (2023) and Gudibande et al. (2023) mention the conjecture that fine-tuning on new factual knowledge may encourage hallucinations. Huang et al. (2023) categorized hallucination causes and formally defined this scenario as capability misalignment. They highlight that limited research addresses capability misalignment due to the challenge of defining the knowledge boundary of LLMs. Kang et al. (2024) showed that when a fine-tuned LLM encounters unknown queries at test time, its responses mimic the responses associated with the unknown examples in the fine-tuning data. Yin et al. (2023) showed that LLMs’ performance is not satisfactory when they face new knowledge in their input contexts and Lee et al. (2023) analyzed the impact of unknown in-context learning examples. To the best of our knowledge, our work is the first to empirically assess the impact of exposure to new knowledge through fine-tuning on tendency of the fine-tuned model to hallucinate.

Quantifying knowledge in LLMs.

SliCK can be seen as a confidence elicitation method for the ground truth label (MM knows (q,a)(q,a) if it is confident that aa is the answer to qq). Existing work derive calibrated confidence from LLMs by examining agreement across multiple samples Kuhn et al. (2023); Manakul et al. (2023); Tian et al. (2023a); Lyu et al. (2024), probing internal representations Azaria and Mitchell (2023); Burns et al. (2022), eliciting verbalized probability Tian et al. (2023b) or direct prompting Kadavath et al. (2022). Kadavath et al. also trained a separate P(IK) model to predict if the LLM knows the answer to qq. The label for P(IK) was approximated by the fraction of correct sampled answers, which is conceptually aligned with PCorrectP_{\mathtt{Correct}} (§3). A key difference is that we also define the SliCK categories, and provide evidence that we capture meaningful and useful categories.

Conclusion

We study the impact of integrating new factual knowledge through fine-tuning on the model’s tendency to hallucinate. We first propose SliCK, a categorization of facts w.r.t. LLM’s knowledge. We then design a controlled study where we isolate the impact of new knowledge and rigorously evaluate its effects. We provide multiple insights on the fine-tuning dynamics, with the following key findings: (1) Acquiring new knowledge via supervised fine-tuning is correlated with hallucinations w.r.t. pre-existing knowledge. (2) LLMs struggle to integrate new knowledge through fine-tuning and mostly learn to use their pre-existing knowledge.

Limitations

Our experiments were conducted using a single LLM, and thus it is unclear whether results will vary with different LLMs. Having said that, our study is extremely compute-heavy and thus challenging to replicate on multiple LLMs: First, our fine-tuning is compute-heavy as its runs are very long as we wanted to analyze the behavior during different stages of fine-tuning (including the overfitting stages). Second, and most importantly, to facilitate our study we needed to annotate a large scale dataset w.r.t the SliCK categories. To derive reliable conclusions, it was crucial to accurately assess the model’s knowledge w.r.t. a single fine-tuning example. In our case we run 170 inference steps per example, i.e., more than 15M15M inference steps to categorize our full dataset.

In addition, since we focus on closed-book QA, the practical implications from our study such as filtering-out Unknown\mathtt{Unknown} fine-tuning examples still require validation in settings involving long-form text generation. To filter-out examples that introduce new factual knowledge in long-form generation tasks, one would need to make adaptations to SliCK and come up with an effective way to compare the sampled answer with the ground-truth to approximate PCorrectP_{\mathtt{Correct}}. We leave this for future work. Long-form generation tasks introduce evaluation challenges, leading to a wide adoption of LLM-based evaluations. Our choice to focus explicitly on closed book QA facilitates more accurate evaluation that enhances the reliability of our findings.

Lastly, we did not test the effect of adding additional fine-tuning examples from diverse tasks into the fine-tuning mixture. While this could more closely approximate a typical instruction fine-tuning scenario, such dataset extension may introduce new factual knowledge in an uncontrollable way, which will limit our findings.

Acknowledgments

We would like to thank Ori Ram, Uri Shaham, Alon Jacovi, Mor Ventura, Yochai Blau, Eyal Ben-David, Avi Caciularu, Avinatan Hassidim and the members of Roi Reichart’s NLP group for reviewing the paper draft and providing valuable feedback. Special thanks to Uri Shaham for assisting in setting up the fine-tuning pipeline during the early stages of our research.

References

Appendix A Data Preprocessing

This section expands §2 with additional details about our data preprocessing steps. The EntityQuestions dataset Sciavolino et al. (2021) consists of train, development and test splits and spans 24 relations. Our train, development and test sets are curated based on the original splits from EntityQuestions. However, we use only 12 relations, since we wanted to reserve some relations for out-of-distribution test set. To avoid cherry-picking, the 12 relations used in our train, development and test sets are randomly sampled. The resulting relations are presented in Tables 3 and 4.

We reserved the remaining 12 relations for out-of-distribution test set. However, we found that in those 12 reserved relations, 5 were too similar to some of the relations that we train on (Table 3), thus we suspected that this could lead to a test set that is not truly out-of-distribution. To address that, we filtered out those relations and were left with 7 relations for our-of-distribution. Specifically we filtered-out the following relations:

P276 was filtered out since it directly overlaps with P131 since for both relations the question in EntityQuestions is of the form “Where is [E] located?”. P276 stands for “location” (https://www.wikidata.org/wiki/Property:P276) and P131 stands for “located in the administrative territorial entity” (https://www.wikidata.org/wiki/Property:P131).

P20, for which the question template is “Where did [E] die?”, was filtered out since it may require knowledge that relates to P19, for which the question template is “Where was [E] born?”. P20 stands for “place of death” (https://www.wikidata.org/wiki/Property:P20) and P19 stands for “place of birth” (https://www.wikidata.org/wiki/Property:P19).

P106, for which the question template is “What kind of work does [E] do?”, was filtered out since it may require knowledge that relates to P800, for which the question template is “What is [E] famous for?”. P106 stands for “occupation” (https://www.wikidata.org/wiki/Property:P106) and P800 stands for “notable work” (https://www.wikidata.org/wiki/Property:P800).

P413, for which the question template is “What position does [E] play?”, was filtered out since it may require knowledge that relates to P800, for which the question template is “What is [E] famous for?”. P413 stands for “position played on team / speciality” (https://www.wikidata.org/wiki/Property:P413) and P800 stands for “notable work” (https://www.wikidata.org/wiki/Property:P800).

P159, for which the question template is “Where is the headquarters of [E]?”, was filtered out since it may require knowledge that relates to P36, for which the question template is “What is the capital of [E]?”. P159 stands for “headquarters location” (https://www.wikidata.org/wiki/Property:P159) and P36 stands for “capital” (https://www.wikidata.org/wiki/Property:P36).

The 7 relations used for out-of-distribution test set are presented in Table 5.

Lastly, we perform two additional filtering steps: (1) To simplify the process of categorizing the examples w.r.t. MM’s knowledge (§3), we filter-out examples with more than 1 correct answer.4.2%4.2\% and 3.9%3.9\% of the EntityQuestions train and test set respectively. (2) We make sure that no subjects or objects overlap between the train and test sets, For example, the subject “Bruce Smith” appears with 2 different relations (P106P106 and P413P413) yielding 2 examples: (“What kind of work does Bruce Smith do?”, “poet”) and (“Where was Bruce Smith born?”, “Faribault”). by filtering-out overlapping examples from the train set.2.1%2.1\% of the EntityQuestions train set.

Appendix B Test performance as Proxy for Hallucinations

We now detail the relation between the test performance in our setting and hallucinations. In our study, poorer performance of a fine-tuned model MD1M_{D1}, compared to another fine-tuned model MD2M_{D2} on the test set, can be attributed to a higher rate of hallucinations in MD1M_{D1}, relative to its pre-existing knowledge, due to the following explanation.

The test set can be conceptually divided into two types of questions. First, there are questions with answers that are unknown to MM. Those questions will remain unknown post fine-tuning, as we make sure that the training set is disjoint from the test set (§A). This means that both MD1M_{D1} and MD2M_{D2} will fail to answer these questions. Thus, the test performance difference between MD1M_{D1} and MD2M_{D2} is mostly attributed to the second type of questions: ones that are known to MM, i.e. MM can answer them correctly since it posses the relevant knowledge. Thus, MD1M_{D1} and MD2M_{D2} must rely on their pre-existing knowledge to answer such questions, and a lower performance on such question can be only categorized as an hallucination w.r.t. pre-existing knowledge.

This section expands §3 with additional details about our PCorrectP_{\mathtt{Correct}} approximation. In our study we approximate PCorrect(q,a;M,T)P_{\mathtt{Correct}}(q,a;M,T) based on the fraction of correct answers to qq sampled from MM. We begin with randomly sampling NexN_{\text{ex}} distinct kk-shot exemplars for each relation in our dataset (§A). Then, to approximate PCorrect(q,a;M,T)P_{\mathtt{Correct}}(q,a;M,T), we use MM to generate answers to qq using each of the NexN_{\text{ex}} exemplars from the relation corresponding to qq. We first use temperature sampling with T=0.5T=0.5 to sample NsampleN_{\text{sample}} answers for each of the NexN_{\text{ex}} exemplars. PCorrect(q,a;M,T>0)P_{\mathtt{Correct}}(q,a;M,T>0) is then approximated by the fraction of correct answers from the total of Nex⋅NsampleN_{\text{ex}}\cdot N_{\text{sample}} predictions. We also generate the greedy decoding prediction (T=0T=0) for each of the NexN_{\text{ex}} exemplars. PCorrect(q,a;M,T=0)P_{\mathtt{Correct}}(q,a;M,T=0) is then approximated by the fraction of correct answers from the total of NexN_{\text{ex}} predictions.Since we can only have one greedy prediction for every k-shot exemplars.

We use k=4k=4 in our study, simply since we found it enough for MM to output answers in the correct format. We use Nex=10N_{\text{ex}}=10 and Nsample=16N_{\text{sample}}=16. The Nsample=16N_{\text{sample}}=16 samples using T=0.5T=0.5 are sampled from Top 40.

The kk exemplars are sampled from the development split. We sample NexN_{\text{ex}} different samples since we found that even when the few-shot exemplars are sampled per-relation, their exact choice still affects the prediction. In §6 and Figure 5 we show evidence that this also improves the quality of our categories.

Below is an example of our 4-shot prompt format, from real example from EntityQuestions with the relation P106P106 representing occupation.https://www.wikidata.org/wiki/Property:P106 The question in this case is “What kind of work does Ron Konopka do?” and the ground truth asnwer is “geneticist”.

To decide whether a sampled answer is correct, we use the Exact Match (EM) metric to compare it with the ground truth answer. The main advantage in this choice is that when EM is True, we know that the answer is correct for 100%100\%. The main potential risk associated with this choice is that we may wrongly classify answers as incorrect due to paraphrases or answers with different granularity levels Wang et al. (2023); Kamalloo et al. (2023); Yona et al. (2024)). To address this, we perform an error analysis on 100 predictions for which EM was False. We randomly sample 50 greedy predictions (T=0T=0) and 50 samples with T=0.5T=0.5. The results are in Table 6. This analysis suggest that in 90%90\% of the cases where EM is False, the predicted answer is indeed incorrect. Which is a reasonable performance for our purpose, especially considering that when EM is True the answer is 100%100\% correct.

Appendix D Data Annotation

we first calculate PCorrect(q,a;M,T=0)P_{\mathtt{Correct}}(q,a;M,T=0) and PCorrect(q,a;M,T>0)P_{\mathtt{Correct}}(q,a;M,T>0) for each (q,a)(q,a) pair in our preprocessed dataset (§2 and §A), using our PCorrect(⋅)P_{\mathtt{Correct}}(\cdot) approximation (§3 and §C). We then use these values to categorize each (q,a)(q,a) pair into one of our four categories (§3 and Figure 2). We provide the full statistics of the categories on the train and test set, as well as the out-of-distribution test set in Tables 3, 4 and 5.

Appendix E Fine-tuning Details

In §4 we examine the effect of new knowledge in the fine-tuning dataset DD on the performance of MDM_{D}, by varying the proportion of Unknown\mathtt{Unknown} examples in DD. When we create variants of DD with exactly X%X\% of Unknown\mathtt{Unknown} and (100−X)%(100-X)\% Known\mathtt{Known} examples, we make sure that the relation distribution remains consistent. To achieve that we sample X%X\% of Unknown\mathtt{Unknown} from each relation.

In §5 we create single-category variants of DD. Since we want to work with a fixed ∣D∣|D| across all variants, we want to make sure that we have ∣D∣|D| examples from each category. To ensure this, we measure the size of the smallest category in each relation (see the “Min” column in Table 3) and define ∣D∣|D| as their sum. In other words, for each relation we calculate the size of the smallest category and sum these values. This leads to ∣D∣=6142|D|=6142, as illustrated by the last column in Table 3. More formally, for each relation r in the training split, and for each category CAT from our 4 SliCK categories, we define CATr\text{CAT}_{\text{r}} to be the examples from category CAT and relation r. Consequently size(CATr)\text{size}(\text{CAT}_{\text{r}}) is the number of the examples in CATr\text{CAT}_{\text{r}}. For example size(\text{size}(HighlyKnown\mathtt{HighlyKnown} P131)=553{}_{\text{P131}})=553 (see Table 3). We then define:

where RTrain\text{R}_{\text{Train}} are the 12 relations from the training set.

Below is an example of our data format in the train, development and test sets, from real example from EntityQuestions with the relation P106P106 representing occupation.https://www.wikidata.org/wiki/Property:P106 The question in this case is “What kind of work does Ron Konopka do?” and the ground truth asnwer is “geneticist”. Answer the following question. What kind of work does Ron Konopka do?

Fine-tuning hypeparameters.

We fine-tune every model for 50 epochs for all our model variants to completely fit the training set, so we can examine all stages of fine-tuning. We use learning rate of 1e-5, a batch size of 128, and a dropout rate of 0.05. We evaluate the models every epoch on the development set. The early_stop stopping criteria is defined to be the epoch with the maximum accuracy on the development set.

Appendix F Train Accuracy on Different 𝙺𝚗𝚘𝚠𝚗𝙺𝚗𝚘𝚠𝚗\mathtt{Known} Categories

In §4.3 we analyze the fine-tuning dynamic and present the training accuracy as function of the fine-tuning duration in Figure 1. For simplicity we treated the Known\mathtt{Known} categories collectively. For reference we also include the plot with the full per-category breakdown in Figure 6.

Appendix G Linear Model

In §4.4 and §4.5 we use a linear model (Equation 1) that predicts that test accuracy and the out-of-distribution test accuracy. We estimate the parameters of this linear model based on results from all our variants of DD used in §4. For all these variants, we measure the test accuracy and the number of Known\mathtt{Known} and Unknown\mathtt{Unknown} fine-tuning examples that MM fits during different fine-tuning stages. This way we collect a dataset with examples of the form (Accuracy,NKn,NUnk)(Accuracy,N_{\text{Kn}},N_{\text{Unk}}), which we use to fit a linear regression model.

Appendix H Out-of-distribution (OOD) Evaluation

In §4.5 we discuss out-of-distribution (OOD) results. In these experiments we simply used our OOD test set consisting of 7 relations unseen during fine-tuning (see §A). When we perform the analysis discussed in §4.1 and §4.2, we additionally evaluated the models on the OOD test set. For completeness, we add here Figure 7, which is the out-of-distribution version of Figure 3. Figure 7(a) presents the OOD test performance as a function of %\% of Unknown\mathtt{Unknown} examples in DD for different fine-tuning duration. The corresponding in-distribution results (Figure 3(a)) were discussed in §4.1. Figure 7(b) presents the OOD test performance for the ablation where we filter-out Unknown\mathtt{Unknown} fine-tuning examples. The corresponding in-distribution results (Figure 3(b)) were discussed in §4.2. We notice that similar trends, just with a smaller overall magnitude of the performance drop, up to 6 points drop compared to up to 14 for in-distribution. This smaller drop magnitude is also reflected in smaller values of ∣βukn∣|\beta_{\text{ukn}}| and ∣βkn∣|\beta_{\text{kn}}| (Table 1).

Appendix I Statistic Significance Tests

In §5 we present Table 2. As mentioned in the caption, we perform statistic significance tests for each column. To this end we compare all the values to the maximal value in this column.

For each subset of the test set, we randomly shuffle all the examples in it, split them up into 100 approximately equally sized subsets, and compute accuracy for each of them for all the models of interest. We then apply paired-sample t-test with p<0.05p<0.05 and p<0.01p<0.01.

In Table 2, the best result is in bold, as well as all the results with statistically non-significant difference from the best with p<0.05p<0.05. We additionally include a copy of Table 2 where all the statistical tests outcomes are annotated, see Table 7. We can see that in almost all cases the difference is statistically significant with p<0.01p<0.01, except two cases where it is only with p<0.05p<0.05 (DNaturalD_{\mathtt{Natural}} Unk\mathtt{Unk} and DMaybeKnownD_{\mathtt{MaybeKnown}} Mkn\mathtt{Mkn}).

Since we also discuss “horizontal” comparisons, where we compare early_stop to Convergence, we additionally run significance tests (not annotated in Table 2) for AllAll, comparing early_stop to Convergence. The difference for DMaybeKnownD_{\mathtt{MaybeKnown}} was not statistically significant while for all others (including DNaturalD_{\mathtt{Natural}}) it was significant with p<0.01p<0.01.

Appendix J The P(True) Case Study

In §6 we used the P(True) metric from Kadavath et al. (2022) as a case study for comparison. In Figure 5 we compare our Unknown\mathtt{Unknown} category vs classifying as Unknown\mathtt{Unknown} based on a threshold of P(True). We calculated P(True) for every (q,a)(q,a) pair in the test set using Kadavath et al. (2022)’s prompt:

We then treated (q,a)(q,a) pairs with P(True) below a threshold as Unknown\mathtt{Unknown}. We experimented with each possible threshold TT in $,accordingtoourtestset.Foreachthreshold, according to our test set. For each thresholdTwethenmeasured(1)howmanyexampleswereclassifiedaswe then measured (1) how many examples were classified as\mathtt{Unknown}outofthetestset,(2)whatwastheaccuracyontheseexamplesafterfine−tuning.WeplottheresultsinFigure5,whereP(True)isrepresentedwiththeyellowlineandourout of the test set, (2) what was the accuracy on these examples after fine-tuning. We plot the results in Figure 5, where P(True) is represented with the yellow line and our\mathtt{Unknown}isrepresentedwiththebluecircle.Asdiscussedin§C,itwasapproximatedusing10defferentsamplesof4−shotexemplars(is represented with the blue circle. As discussed in §C, it was approximated using 10 defferent samples of 4-shot exemplars (N_{\text{ex}}=10).Wealsochecksmallervaluesof). We also check smaller values ofN_{\text{ex}}andplottheresultswiththeblueline.Theaccuracyafterfine−tuningforalltheresultsismeasuredafterfine−tuningwithand plot the results with the blue line. The accuracy after fine-tuning for all the results is measured after fine-tuning withD_{\mathtt{Natural}}$ (§5).

In this work we showed that fitting Unknown\mathtt{Unknown} fine-tuning examples negatively affects the test performance. However, this negative effect manifests as a form of overfitting. From practical perspective, we showed that we can mitigate overfitting by either using early-stopping or filtering-out Unknown\mathtt{Unknown} examples from the fine-tuning dataset.

We now perform a preliminary experiment where check whether fine-tuning the model to abstain from Unknown\mathtt{Unknown} examples can also be a potential mitigation. Specifically, we replace the label of the Unknown\mathtt{Unknown} fine-tuning examples with the expression “I don’t know” and test whether this mitigates the observed overfitting.

Table 8 presents the %\% of the test questions that were answered (i.e. MDM_{D} did not respond with “I don’t know”) and the accuracy on those questions. This experiment was conducted on the DD variant with 50%50\% Unknown\mathtt{Unknown}. The first row is for the original result with DD as a reference and the second row is for the results with DIDKD_{\mathtt{IDK}}, where the ground-truth label of the 50%50\% of the Unknown\mathtt{Unknown} examples in DD was replaced with “I don’t know”

Consistent with the findings from previous work Zhang et al. (2023), we observe an improved accuracy on willingly answered test examples (when comparing DD vs DIDKD_{\mathtt{IDK}}). When we compare early_stop vs Convergence for DD we observe a performance drop (43.0→38.843.0\rightarrow 38.8) which illustrates the overfitting effect. However, we observe that re-labeling the Unknown\mathtt{Unknown} examples with uncertainty expression seem to reduce the risk of overfitting. Specifically, the accuracy for DIDKD_{\mathtt{IDK}} remains 61.861.8 for both early_stop and Convergence, with a small decrease on the number of willingly answered questions (58.7→55.658.7\rightarrow 55.6)