Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge Conflicts

Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, Yu Su

Introduction

After pre-training on massive corpora, large language models (LLMs) (Brown et al., 2020; Chowdhery et al., 2022; Ouyang et al., 2022; OpenAI, 2022; 2023; Zeng et al., 2023; Touvron et al., 2023a) have formed a wealth of parametric memory, such as commonsense and factual knowledge (Petroni et al., 2019; Li et al., 2022; Zhao et al., 2023). However, such parametric memory may be inaccurate or become outdated (Liska et al., 2022; Luu et al., 2022) due to misinformation in the pre-training corpus or the static nature of parametric memory, known to be a major cause for hallucinations (Elazar et al., 2021; Shuster et al., 2021; Ji et al., 2023).

ToolIn the rest of the paper we use “tool-augmented LLMs” because retrievers are one type of tools, but tools are not limited to retrievers (consider, e.g., a question answering tool). (Schick et al., 2023; Qin et al., 2023) or retrieval augmentation (Mallen et al., 2022; Shi et al., 2023; Ram et al., 2023) has emerged as a promising solution by providing external information as new evidence to LLMs, such as ChatGPT Plugins and New Bing. However, external evidence, inevitably, could conflict with LLMs’ parametric memory. We refer to external evidence that conflicts with parametric memory as counter-memory. In this paper, we seek to answer the question: how receptive are LLMs to external evidence, especially counter-memory? A solid understanding of this question is an essential stepping stone for wider application of tool-augmented LLMs. Not only does this relate to overcoming the limitations of LLM’s static parametric memory, but it is also associated with direct safety concerns. For example, what if a third-party tool, either by the developer or hijacked by attackers, intentionally returns disinformation? Will LLMs be deceived?

We present the first comprehensive and controlled investigation into the behavior of LLMs when encountering counter-memory. A key challenge lies in how to construct the counter-memory. Prior work employs various heuristics, such as negation injection (Niu & Bansal, 2018; Kassner et al., 2021; Gubelmann & Handschuh, 2022) and entity substitution (Longpre et al., 2021; Zhou et al., 2023), and finds that language models (both large and small) tend to be stubborn and cling to their parametric memory. However, such heuristic word-level editing results in incoherent counter-memory (see an example in Section 4.1), which may make it trivial for LLMs to detect and thus neglect the constructed counter-memory. It is unclear how the prior conclusions translate to real-world scenarios, where counter-memory is more coherent and convincing.

We propose a systematic framework to elicit the parametric memory of LLMs and construct the corresponding counter-memory. We design a series of checks, such as entailment from parametric memory to the answer, to ensure that the elicited parametric memory is indeed the LLM’s internal belief. For the counter-memory, instead of heuristically editing the parametric memory, we instruct an LLM to directly generate a coherent passage that factually conflicts with the parametric memory. After obtaining a large pool of parametric memory and counter-memory pairs, we then comprehensively examine LLMs’ behavior in different knowledge conflict scenarios, including 1) when only counter-memory is present as external evidence and 2) when both parametric memory and counter-memory are present.

Our investigation leads to a series of interesting new findings. We highlight the following:

LLMs are highly receptive to external evidence if that is the only evidence, even when it conflicts with their parametric memory. This contradicts the prior wisdom (Longpre et al., 2021), and we attribute this to the more coherent and convincing counter-memory constructed through our framework. On the other hand, this also suggests that LLMs may be easily deceived by, e.g., disinformation from malicious (third-party) tools.

However, with both supportive and contradictory evidence to their parametric memory, LLMs show a strong confirmation bias (Nickerson, 1998) and tend to cling to their parametric memory. This reveals a potential challenge for LLMs to unbiasedly orchestrate multiple pieces of conflicting evidence, a common situation encountered by generative search engines.

Related Work

After pre-training, language models have internalized a vast amount of knowledge into their parameters (Roberts et al., 2020; Jiang et al., 2020), also known as parametric memory. Many past studies have explored the elicitation of parametric memory in language models, such as commonsense or factual knowledge probing (Petroni et al., 2019; Lin et al., 2020; Zhang et al., 2021; West et al., 2022). Such parametric memory could help solve downstream tasks (Wang et al., 2021; Yu et al., 2023; Sun et al., 2023). However, previous work has discovered that language models only memorize a small portion of the knowledge they have been exposed to during pre-training (Carlini et al., 2021; 2023) due to model’s limited memorization abilities. In addition, the parametric memory may become outdated (Lazaridou et al., 2021; De Cao et al., 2021). Such incorrect and outdated parametric memory may show as hallucinations (Elazar et al., 2021; Shuster et al., 2021; Ji et al., 2023). Although some methods are proposed to edit knowledge in language models (Dai et al., 2022; Meng et al., 2022; 2023), they typically require additional modifications on model weights without evaluating the consequences on models’ other aspects such as performances and are limited to factual knowledge.

Tool-augmented Language Models

To address the limitations of parametric memory, external tools such as retrievers are used to augment language models with up-to-date information, namely tool-augmented (Nakano et al., 2021; Yao et al., 2023; Qin et al., 2023; Schick et al., 2023; Lu et al., 2023) or retrieval-augmented (Guu et al., 2020; Khandelwal et al., 2020; Izacard & Grave, 2021; Borgeaud et al., 2022; Zhong et al., 2022) language models. Such a framework, which has proven its efficacy in enhancing large language models (Shi et al., 2023; Ram et al., 2023; Mallen et al., 2022), is adopted in real-world applications such as New Bing and ChatGPT Plugins. Inevitably, the external evidence could conflict with the parametric memory. However, the behavior of LLMs in knowledge conflict scenarios remains under-explored, and unraveling it holds significance for wider applications of tool-augmented LLMs.

Knowledge Conflict

To perform controlled experiments, knowledge conflict is often simulated with counter-memory constructed upon parametric memory. Heuristic counter-memory construction methods such as negation injection (Niu & Bansal, 2018; Kassner et al., 2021; Petroni et al., 2020; Pan et al., 2021) have been developed. Furthermore, entity substitution (Longpre et al., 2021; Chen et al., 2022; Si et al., 2023; Zhou et al., 2023) replaces all mentions of the answer entity in parametric memory with other entities to construct counter-memory. However, these methods are limited to word-level editing, leading to low overall coherence in the counter-memory. We instead instruct LLMs to generate counter-memory from scratch to ensure high coherence.

Experimental Setup

In this section, we describe our framework for eliciting high-quality parametric memory from LLMs and constructing the corresponding counter-memory, as well as the evaluation metrics.

Following prior work (Longpre et al., 2021; Chen et al., 2022), we adopt question answering (QA) task as the testbed for knowledge conflict experiments. In addition to an entity-based QA dataset (PopQA), we include a multi-step reasoning dataset (StrategyQA) for diversifying the questions studied in the experiments. Specifically,

PopQA (Mallen et al., 2022) is an entity-centric QA dataset that contains 14K questions. Data for PopQA originates from triples in Wikidata. Employing custom templates tailored to relationship types, the authors construct questions through the substitution of the subject within knowledge triples. PopQA defines the popularity of a question based on the monthly Wikipedia page views associated with the entity mentioned in the question.

StrategyQA (Geva et al., 2021) is a multi-step fact reasoning benchmark that necessitates the implicit question decomposition into reasoning steps. The questions are built around Wikipedia terms and cover a wide range of strategies, which demand the model’s capability to select and integrate relevant knowledge effectively. The language model is expected to provide a True or False answer.

2 Parametric Memory Elicitation

Step 1 in Figure 1 illustrates how we elicit parametric memory: in a closed-book QA fashion, LLMs recall their parametric memory to answer questions without any external evidence. Specifically, given a question, e.g., “Who is the chief scientist of Google DeepMind”, LLMs are instructed to provide an answer “Demis Hassabis” and its supporting background information about how Demis founded and led DeepMind in detail. We cast the detailed background as parametric memory because the answer only represents the conclusion of parametric memory w.r.t. the given question.

Table 1 shows the closed-book results of LLMs on PopQA and StrategyQA. Notably, LLMs may respond with “Unknown” when no evidence is provided in the context, particularly in ChatGPT. Such answer abstention (Rajpurkar et al., 2018) suggests that LLMs fail to recall valid memory associated with the given question, so we discard them. For comprehensiveness, we also keep the examples that LLMs answer incorrectly in the closed-book paradigm because the wrong answer and associated memory are also stored in model parameters.

3 Counter-memory Construction

As depicted in Figure 1, at Step 2, we reframe the memory answer “Demis Hassabis” to a counter-answer (e.g., “Jeff Dean”). Concretely, for PopQA, we substitute the entity in the memory answer with a same-type entity (e.g., from Demis to Jeff); while in StrategyQA, we flip the memory answer (e.g., from positive sentence to negative sentence). With counter-answer “Jeff Dean”, we instruct ChatGPT We leverage ChatGPT for its cost-effectiveness and its on-par counter-memory generation ability with GPT-4. In our pilot study (based on 1000 instances), LLMs showed the same level of receptiveness to counter-memory generated by both ChatGPT and GPT-4. to make up supporting evidence that Jeff Dean serves as chief scientist of DeepMind. We term such evidence that conflicts with parametric memory as counter-memory.

Since the counter-memory is generated from scratch by powerful generative LLMs, it is more coherent compared to previous word-level editing methods (Longpre et al., 2021; Chen et al., 2022) performed on parametric memory. Both generated parametric memory and counter-memory could serve as external evidence for later experiments on LLMs in knowledge conflicts. Please refer to Appendix B.1 for more details of evidence construction in each dataset.

4 Answer-evidence Entailment Checking

An ideal piece of evidence should strongly support its answer. For instance, the parametric memory about Demis and DeepMind should clearly support the corresponding memory answer that Demis is the chief scientist of DeepMind. Similarly, counter-memory should clearly support the corresponding counter-answer as well. Therefore, for Step 3 shown in Figure 1, we utilize a natural language inference (NLI) model for support-checking to ensure the evidence indeed entails the answer. Specifically, we use the state-of-the-art NLI model DeBERTa-V2 (He et al., 2021)https://huggingface.co/microsoft/deberta-v2-xxlarge-mnli. to determine whether both the parametric memory and counter-memory support their corresponding answers. We only keep the examples where both answers are supported for subsequent experiments.

To ensure the reliability of the selected NLI model, we manually evaluated 200 random examples. The model achieves a 99% accuracy, demonstrating its effectiveness for our purpose.

5 Memory Answer Consistency

We adopt another check (Step 4 of Figure 1) for further ensuring the data quality. If the parametric memory we elicit is truly the internal belief of an LLM’s, presenting it explicitly as evidence should lead the LLM to provide the same answer as in the closed-book setting (Step 1). Therefore, in the evidence-based QA task format, we use the parametric memory as the sole evidence and instruct LLMs to answer the same question again. For example, given the parametric memory about Demis and DeepMind, LLMs should have a consistent response with the previous memory answer, that Demis is the chief scientist of DeepMind.

However, the answer inconsistency results in Table 4 show that LLMs may still change their answers when the parametric memory obtained in Step 1 is explicitly presented as evidence. This suggests that the LLM’s internal belief on this parametric memory may not be firm (e.g., there may competing answers that are equally plausible based on the LLM). We filter out such examples to ensure the remaining ones well capture an LLM’s firm parametric memory.

After undergoing entailment and answer consistency checks, the remaining examples are likely to represent firm parametric memory and high-quality counter-memory, which lay a solid foundation for subsequent knowledge conflict experiments. Some examples from the final PopQA data are shown in Table 2 and the statistics of the final datasets are shown in Table 4. Please refer to Appendix B.2 for more details for Step 3 and 4 and examples.

6 Evaluation Metrics

A single generation from an LLM could contain both the memory answer and the counter-answer, which poses a challenge to automatically determine the exact answer from an LLM. To address this issue, we transform the free-form QA to a multiple-choice QA format by providing a few options as possible answers. This limits the generation space and helps determine the answer provided by LLMs with certainty. Specifically, for each question from both datasets, LLMs are instructed to select one answer from memory answer (Mem-Ans.), counter-answer (Ctr-Ans.), and “Uncertain”. Additionally, to quantify the frequency of LLMs sticking to their parametric memory, we adopt the memorization ratio metric (Longpre et al., 2021; Chen et al., 2022):

where fmf_{m} is the frequency of memory answer and fcf_{c} is that of counter-answer. Higher memorization ratios signify LLMs relying more on their parametric memory, while lower ratios indicate more frequent adoption of the counter-memory.

Experiments

We experiment with LLMs in the single-source evidence setting where counter-memory is the sole evidence presented to LLMs. Such knowledge conflict happens when LLMs are augmented with tools returning single external evidence such as Wikipedia API (Yao et al., 2023). In particular, for counter-memory construction, we would apply 1) the entity substitution counter-memory method, a widely-applied strategy in previous work, and 2) our generation-based method.

Following previous work (Longpre et al., 2021; Chen et al., 2022), we substitute the exactly matched ground truth entity mentions in the parametric memory with a random entity of the same type. The counter-memory is then used as the sole evidence for LLMs to answer the question. Here is an example:

Evidence: Washington D.C. London, USA’s capital, has the Washington Monument.

Question: What is the capital city of USA? Answer by ChatGPT: Washington D.C.

Figure 2 shows the results with this approach. Observably, although the instruction clearly guides LLMs to answer questions based on the given counter-memory, LLMs still stick to their parametric memory instead, especially for three closed-sourced LLMs (ChatGPT, GPT-4, and PaLM2). This observation is aligned with previous work (Longpre et al., 2021). The reasons may stem from the incoherence of the evidence built with substitution: In the given example, although “Washington D.C.” is successfully substituted by “London”, the context containing Washington Monument and USA still highly correlate with the original entity, impeding LLMs to generate London as the answer. Furthermore, when comparing Llama2-7B and Vicuna-7B to their larger counterparts in the same series (i.e., Llama2-70B and Vicuna-33B), we observe that the larger LLMs are more inclined to insist on their parametric memory. We suppose that larger LLMs, due to their enhanced memorization and reasoning capabilities, are more sensitive to incoherent sentences.

LLMs are highly receptive to generated coherent counter-memory.

To alleviate the incoherence issue of the above counter-memory, we instruct LLMs to directly generate coherent counter-memory following the steps aforementioned (Figure 1). Figure 2 shows the experimental results with generation-based counter-memory, from which we can have the following observations:

First, LLMs are actually highly receptive to external evidence if it is presented in a coherent way, even though it conflicts with their parametric memory. This contradicts the prior conclusion (Longpre et al., 2021) and the observation with entity substitution counter-memory shown in Figure 2. Such high receptiveness in turn shows that the counter-memory constructed through our framework is indeed more coherent and convincing. We manually check 50 stubborn (i.e., “Mem-Ans.”) cases and find that most of them are due to hard-to-override commonsense or lack of strong direct conflicts. Detailed analyses can be found in Appendix B.3.

Second, many of the generated counter-memory are disinformation that misleads LLMs to the wrong answer. Concerningly, LLMs appear to be susceptible to and can be easily deceived by such disinformation. Exploring methods to prevent LLMs from such attacks when using external tools warrants significant attention in future research.

Third, the effectiveness of our generated counter-memory also shows that LLMs can generate convincing dis- or misinformation, sufficient to mislead even themselves. This raises concerns about the potential misuse of LLMs.

2 Multi-source Evidence

Multi-source evidence is a setting where multiple pieces of evidence that either supports or conflicts with the parametric memory are presented to LLMs. Such knowledge conflicts can happen frequently, e.g., when LLMs are augmented with search engines having diverse or even web-scale information sources. We study the evidence preference of LLMs from different aspects of evidence, including popularity, order, and quantity. By default, the order of evidence is randomized in all experiments in Section 4.2, if not specified otherwise.

Step 5 in Figure 1 illustrates how we instruct LLMs to answer questions when both parametric memory and counter-memory are presented as evidence. Figure 3 shows the memorization ratio of different LLMs w.r.t. the question popularity on PopQA.

First, compared with when only the generated counter-memory is presented as evidence (single-source), both LLMs demonstrate significantly higher memorization ratios when parametric memory is also provided as evidence (multi-source), especially in the case of GPT-4. In other words, when faced with conflicting evidence, LLMs often prefer the evidence consistent with their internal belief (parametric memory) over the conflicting evidence (counter-memory), demonstrating a strong confirmation bias (Nickerson, 1998). Such properties could hinder the unbiased use of external evidence in tool-augmented LLMs.

Second, for questions regarding more popular entities, LLMs demonstrate a stronger confirmation bias. In particular, GPT-4 shows an 80% memorization ratio for the most popular questions. This may suggest that LLMs form a stronger belief in facts concerning more popular entities, possibly because they have seen these facts and entities more often during pre-training, which leads to a stronger confirmation bias.

LLMs demonstrate a noticeable sensitivity to the evidence order.

Previous work has shown a tendency in tool-augmented language models to select evidence presented in the top place (BehnamGhader et al., 2022) and the order sensitivity in LLMs (Lu et al., 2022). To demystify the impact of the evidence-presenting order in LLMs, we respectively put parametric memory and counter-memory as the first evidence in multi-source settings. As a reference, the results of first evidence randomly selected from the two are also reported in Table 4.2. In line with the popularity experiment, we use the same LLMs.

We observe that, with the exception of GPT-4, other models demonstrated pronounced order sensitivity, with fluctuations exceeding 5%. It’s especially concerning that the variations in PaLM2 and Llama2-7B surpassed 30%. When evidence is presented first, ChatGPT tends to favor it; however, PaLM2 and Llama2-7B lean towards later pieces of evidence. Such order sensitivity for evidence in the context may not be a desirable property for tool-augmented LLMs. By default, the order of evidence is randomized in other experiments in this section.

LLMs follow the herd and choose the side with more evidence.

In addition to LLM-generated evidence (parametric memory and counter-memory), we also extend to human-crafted ones such as Wikipedia. These highly credible and accessible human-written texts are likely to be retrieved as evidence by real-world search engine tools. We adopt Wikipedia passages from PopQA and manually annotated facts from StrategyQA with post-processing to ensure that the ground truth answer can indeed be deduced. Please refer to Appendix B.4 for more processing details.

To balance the quantity of evidence supporting memory answer and counter-answer, we create additional evidence through the method mentioned in Section 3.3, with the goal of achieving a balanced 2:2 split at most between parametric memory and counter-memory evidence. Table 6 shows the memorization ratio under different proportions between parametric memory-aligned evidence and counter-memory. We have three main observations: 1) LLMs generally provide answers backed by the majority of evidence. The higher the proportion of evidence supporting a particular answer, the more likely LLMs will return that answer. 2) The confirmation bias becomes increasingly obvious with a rise in the quantity of parametric memory evidence, despite maintaining a consistent relative proportion (e.g., \nicefrac12\nicefrac{{1}}{{2}} vs. \nicefrac24\nicefrac{{2}}{{4}}). 3) Compared to other LLMs, GPT-4 and Vicuna-33B are less receptive to counter-memory across all proportions of evidence. Particularly, regardless of more pieces of evidence supporting the counter-answer (ratio \nicefrac13\nicefrac{{1}}{{3}}), these two models still noticeably cling to their parametric memory. These observations once again signify the confirmation bias in LLMs.

LLMs can be distracted by irrelevant evidences.

We further experiment on more complicated knowledge conflict scenario. We are interested in this question: Tools such as search engine may return irrelevant evidence — What if irrelevant evidence is presented to LLMs? When irrelevant evidence is presented, LLMs are expected to 1) abstain if no evidence clearly supports any answer and 2) ignore irrelevant evidence and answer based on the relevant ones. To set up, we regard top-ranked irrelevant passages retrieved by Sentence-BERT embeddingshttps://huggingface.co/sentence-transformers/multi-qa-mpnet-base-dot-v1. (Reimers & Gurevych, 2019) as irrelevant evidence (i.e., sentences unrelated to the entities shown in the question). The experimental results on PopQA are presented in Table 7. We find that: 1) With only irrelevant evidence provided, LLMs can be distracted by them, delivering irrelevant answers. And this issue is particularly concerning in Llama2-7B. Meanwhile, as more irrelevant evidence is introduced, LLMs become less likely to answer based on their parametric memory. 2) With both relevant and irrelevant evidence provided, LLMs can filter out the irrelevant ones to a certain extent. However, as the quantity of irrelevant evidence increases, such an ability diminishes, especially Llama2-7B.

Conclusion

In this work, we propose a systematic framework to elicit the parametric memory of LLMs, construct counterpart counter-memory, and design a series of checks to entire their quality. With these parametric memory and counter-memory as external evidence, we simulate comprehensive scenarios as controlled experiments to unravel the behaviors of LLMs in knowledge conflicts. We find that LLMs are highly receptive to counter-memory when it is the only evidence presented in a coherent way. However, LLMs also demonstrate a strong confirmation bias toward parametric memory when both supportive and contradictory evidence to their parametric memory are present. In addition, we show that LLMs’ evidence preference is influenced by the popularity, order, and quantity of evidence, none of which may be a desired property for tool-augmented LLMs. Finally, the effectiveness of our framework also demonstrates that LLMs can generate convincing misinformation, which poses potential ethical risks. We hope our work provides a solid evaluation testbed and useful insights for understanding, improving, and deploying tool-augmented LLMs in the future.

Ethics Statement

Our study highlights a serious concern: LLMs can be instructed to make up coherent and convincing fake information. This underscores the potential misuse of these models if left unchecked. As researchers, it is our duty to address this pressing issue. The risks associated with the misuse of LLMs demand robust safeguards and prevention measures, requiring concerted effort from the wider research community. To this end, we commit to careful distribution of the data generated through our research, ensuring it serves strictly for research purposes. Our goal is to mitigate the risks while maximizing the benefits offered by LLMs.

Reproducibility Statement

Our experiments utilize three closed-sourced LLMs accessed via API, as well as five open-sourced LLMs. We have increased reproducibility by including the prompts used in our experiments in Appendix C. As for the versions of the closed-sourced LLMs, we used ChatGPT-0301, GPT-4-0314, and Chat-Bison-001 of PaLM2 in all our tests.

References

Appendix

Within this supplementary material, we elaborate on the following aspects:

Appendix A Additional Discussion

As a proxy of convincing degree, the length of evidence may affect the preference of LLMs. To verify it, we categorize the examples based on the length ratio between parametric memory and counter-memory, i.e., <0.8<0.8, >1.2>1.2, and [0.8,1.2][0.8,1.2], which are distinguishable in the data samples.Consistent results and observations are found in results with other splits. Figure A.2 shows the answer distribution within each category. It is evident that ChatGPT tends to adopt the longer side, especially in StrategyQA, where longer evidence generally indicates more reasoning steps.

To explore the largest impact of evidence length, we further explore the scenarios with extremely short evidence. Specifically, we present the answer as evidence to LLMs directly and investigate whether they adopt such a short evidence without any concrete explanations. We alternately replace either parametric memory or counter-memory with their respective supporting answers, while keeping the other one intact. This results in memory answer vs. counter-memory and counter-answer vs. parametric memory. Table A.1 shows the results of PopQA: shorter counter-memory evidence (counter-answer) is less likely to be considered by LLMs (56.7% to 18.8%). However, shortening parametric memory evidence into memory answer does not affect the preferences of LLMs much; interestingly, it is even more favored by LLMs (42.7% to 43.9%). In other words, persuading LLMs to embrace counter-memory needs informative and solid evidence. In contrast, short evidence that aligns with parametric memory is acceptable enough by LLMs as the associated memory is encoded in the parameters already. This observation indicates the parametric memory we elicit could well be the firm beliefs of LLMs. More importantly, this unequal receptiveness to evidence further highlights the presence of strong confirmation bias in LLMs, a potentially significant limitation when they are used in tool-augmented applications.

LLMs demonstrate a deficiency in information integration.

In real-world scenarios, a complex query may require fragmented evidence gathered from different sources to have the final answer. As a multi-step reasoning dataset, StrategyQA provides multiple separate pieces of evidence related to sub-questions. Therefore, we take StrategyQA as an ideal sample dataset for such exploration. In the standard mode, we merge these facts to construct an intact piece of evidence. However, in this setting, we treat each fact as an individual piece of evidence, without any consolidation. The results in Figure A.2 clearly show: after the original evidence (parametric memory or counter-memory) used by ChatGPT is fragmented, ChatGPT shifts to consider the other intact evidence (counter-memory or parametric memory) in 38.2% examples, indicating the limited abilities of LLMs to integrate fragments of evidence. This observation also suggests that the same external evidence in different formats (fragmented or whole) may have different effects on LLMs in the tool-augmented systems. Therefore, from the perspective of external tools, it is worth exploring the presentation of evidence in an easy-to-use format for LLMs in the future.

Appendix B Experimental Setup Details

To construct high-quality counter-memory, we incorporate ChatGPT as a generator to produce text at a human-written level. Specifically, we first reframe the memory answer to construct the counter-answer. For different datasets, we utilize different strategies.

Due to the PopQA is a entity-centric QA dataset, we adopt the following principles: (i) If the memory answer is right, we directly adopt the triplets provided by PopQA. (ii) If the memory answer is wrong, we substitute the object entities in the triplets with those of the same relation from the ground truth. Filters are applied based on exact matching to prevent any overlap between the selected entities and the candidate ground truth. Subsequently, we use a template to generate claims in a natural language format based on the triplets.

Considering that the output of StrategyQA is “True” or “False”, it cannot be directly used as a claim. Therefore, we employ ChatGPT to generate two claims corresponding to “True” and “False”, respectively. Based on the output, the generated claims are dynamically classified as memory answer and counter-answer. To ensure high-quality and control format, we adopt the in-context learning strategy and use three demonstrations.

After obtaining the counter-answer, we instruct the ChatGPT to generate the counter-memory.

B.2 Dataset Details

The dataset scale at each step are presented in the Table B.3. We also report the inconsistency type distribution in Table B.4. And some examples of answer inconsistency on LLMs are presented in Table B.5. In Table B.6, we show more examples in the final datasets.

B.3 Examples of Stubbornness in Response to Parametric Memory

In Table B.7, we present some examples which are stubborn to give memory answer even only the counter-memory evidence given. Upon manually scrutinizing 50 randomly selected samples, we discover that ambiguity in counter-memory, commonsense question leading to unacceptable counter-memory, or highly suggestive questions, account for 34 of these instances. This implies that only a minimal fraction of LLMs demonstrate stubbornness towards parametric memory, reaffirming that LLMs maintain open in the single source setting.

B.4 Process for Human-written Evidence

Despite the availability of retrieved Wikipedia passages in the PopQA dataset, not all questions have a high-quality inferential passage (i.e., containing the ground truth). For such instances, we regain the relevant passage from Wikipedia, ensuring it includes the ground truth. However, a small portion of data (around 400 instances) lack inferential passages even on Wikipedia. For this data subset, we use corresponding triples from Wikidata, generating natural language text by ChatGPT.

As for StrategyQA, the facts in it are manually written, ensuring each fact supports the ground truth, and therefore require no additional modifications.

B.5 Irrelevant Evidence

We collect irrelevant evidence for the question from the human-written corpus (i.e., Wikipedia passages provided by PopQA). Specifically, we use SentenceBERT to retrieve the top 3 sentences with the highest similarity to the question. We limit our search to data within the same question type. Note that we exclude any evidence that includes the entity mentioned in the parametric memory or counter-memory , as it would affect the arrangement of our options. The method for constructing options for irrelevant evidence is based on the template provided in the Table B.2.

B.6 Fragmented Evidence

The StrategyQA dataset incorporates human-written facts associated with each sub-question. In the standard mode, we merge these facts to construct an intact piece of evidence. However, in Section A, we treat each fact as an individual piece of evidence, without any consolidation.

Appendix C Prompts List

In Table C.8, we provide a comprehensive list of all the prompts that have been utilized in this study, offering a clear reference for understanding our experimental approach.