DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, Pengcheng He
Introduction
Large language models (LLMs) have demonstrated significant potential in numerous natural language processing (NLP) applications (Brown et al., 2020; OpenAI, 2022; 2023). However, despite the continued increase in performance and the emergence of new capabilities from scaling LLMs (Wei et al., 2022a), their tendency to “hallucinate”, i.e., generate content that deviates from real-world facts observed during pretraining (Ji et al., 2023), remains a persistent challenge. This represents a significant bottleneck in their deployment especially for high-stakes applications (e.g., clinical/legal settings) where reliable generation of trustworthy text is crucial.
While the exact reasons for LMs’ hallucinations are not completely understood, a possible reason is due to the maximum likelihood language modeling objective which minimize the forward KL divergence between the data and model distributions. This objective potentially results in a model with mass-seeking behavior which causes the LM to assign non-zero probability to sentences that are not fully consistent with knowledge embedded in the training data. Empirically, an LM trained with the next-word prediction objective on finite data has been shown to result in a model that use linguistic knowledge to recognize the superficial patterns in the training examples, instead of recognizing and generating the real-world facts extracted from the training corpus (Ji et al., 2023).
From a model interpretability perspective, transformer LMs have been loosely shown to encode “lower-level” information (e.g., part-of-speech tags) in the earlier layers, and more “semantic” information in the later layers (Tenney et al., 2019). More recently, Dai et al. (2022) find that “knowledge neurons” are distributed in the topmost layers of the pretrained BERT model. Meng et al. (2022) show that factual knowledge can even be edited by manipulating a specific set of feedforward layers within an autoregressive transformer LM. We propose to exploit this modular encoding of knowledge to amplify the factual knowledge in an LM through a contrastive decoding approach, where the output probability over the next word is obtained from the difference in logits obtained from a higher layer versus a lower layer. By emphasizing the knowledge from higher layers and downplaying the lower or intermediate layer knowledge, we can potentially make LMs more factual and consequently reduce hallucinations.
An illustration of this idea for a simple example is shown in Figure 1. While “Seattle” maintains high probability throughout all the layers—presumably because it is a syntactically plausible answer—the probability of the true answer “Olympia” increases after the higher layers inject more factual knowledge. Contrasting the differences between the different layers can thus reveal the true answer in this case. Based on this concept, we propose a new decoding method, Decoding by Contrasting Layers (DoLa), for better surfacing factual knowledge embedded in an LLM without retrieving external knowledge or additional fine-tuning.
Experiments on TruthfulQA (Lin et al., 2022) and FACTOR Muhlgay et al. (2023) demonstrate that DoLa is able to increase the truthfulness of the models of the LLaMA family (Touvron et al., 2023). Further experiments on chain-of-thought reasoning for StrategyQA (Geva et al., 2021) and GSM8K (Cobbe et al., 2021) also show that it can facilitate more factual reasoning. Finally, experiments on open-ended text generation results (evaluated with GPT-4) show that when compared with the original decoding method, DoLa can generate informative and significantly more factual responses that lead to better ratings. From an efficiency perspective, we find that DoLa causes only a small additional latency in the decoding process, suggesting it as a practical and useful decoding strategy for improving the truthfulness of LLMs.
Method
Recent language models are consists of an embedding layer, stacked transformer layers, and an affine layer for predicting the next-word distribtution. Given a sequence of tokens , the embedding layer first embeds the tokens into a sequence of vectors . Then would be processed by each of the transformer layers successively. We denote the output of the -th layer as . Then, the vocabulary head predicts the probability of the next token
where is the vocabulary set.
Instead of applying just on the final layer, our approach contrasts the higher-layer and lower-layer information to obtain the probability of next token. More specifically, for the lower layers, we also compute the probability of the next tokens using ,
The idea of applying language heads directly to the hidden states of the middle layers, known as early exit (Teerapittayanon et al., 2016; Elbayad et al., 2020; Schuster et al., 2022), has proven to be an effective inference method even without special training process (Kao et al., 2020), as the residual connections (He et al., 2016) in transformer layers make the hidden representations gradually evolve without abrupt changes. Using to represent for notational brevity, we then compute the probability of the next token by,
Here, layer is referred to as the premature layer, while the final layer is referred to as the mature layer. The operator , to be elaborated further in Section 2.3, is used to contrast between the output distributions from the premature layer and the mature layer by computing the difference between two distributions in the log domain. The premature layer is dynamically selected in each decoding step using a distributional distance measure (we use the Jensen-Shannon Divergence) between the mature layer and all the candidate layers in . We discuss in more detail in Section 2.1 and Section 2.2. The motivation for selecting the layer with the highest distance as the premature layer is to maximize the difference between the mature/premature layers.
We conduct preliminary analysis with the 32-layer LLaMA-7B (Touvron et al., 2023) model to motivate our approach. Here, we compute the Jensen-Shannon Divergence (JSD) between the early exiting output distributions and the final layer output distribution , to show how the early exiting outputs are different from the final layer outputs. Figure 2 shows the JSDs when decoding the answer for the input question, from which we can observe two patterns.
The first type of pattern is when predicting important name entities or dates, such as Wole Soyinka and 1986 in Figure 2, which require factual knowledge. We observe the calculated JSD would be still extremely high in the higher layers. This pattern indicates that the model is still changing its predictions in the last few layers, and potentially injecting more factual knowledge into the predictions.
The second type of pattern is when predicting function words, such as was, the, to, in, and the tokens that are copied from the input question, such as first Nigerian, Nobel Prize. When predicting these “easy” tokens, we can observe that the JSD becomes very small from the middle of the layers. This finding indicates that the model has already decided what token to generate in the early layers, so it just keeps the output distribution almost unchanged in the higher layers. This finding is also consistent with the assumptions in early exiting language models (Schuster et al., 2022).
Qualitatively, when the next-word prediction requires factual knowledge, LLaMA seems to to change the predictions in the higher layers. Contrasting the layers before/after a sudden change may therefore amplify the knowledge emerging from the higher layers and make the model more rely more on its factual internal knowledge. Moreover, this evolution of information seems to vary token by token. In our proposed method, we need to accurately select the premature layer that contains plausible but less factual information, which may not always stay in the same early layer. We propose an approach for dynamic premature later selection as illustrated in Figure 3.
2 Dynamic Premature Layer Selection
To magnify the effective of contrastive decoding, the optimal premature layer to select should ideally be the layer that is the most different from the final-layer outputs. To allow for dynamic premature layer selection at each time step, we adopt the following measure of distance between the next-word distributions obtained from two layers,
where is the Jensen-Shannon divergence. The premature layer, i.e., the -th layer (), is then selected as the layer with the maximum divergence among the subset of early layers,
where is the set of candidate early layers considered for premature layer selection. For LLaMA models with a varying number of layers, we divide the transformer layers into 2 to 4 buckets of based on their total number of layers, in order to focus on contrasting from a certain range of layers. We still use a validation set to select the best bucket depending on the task at hand. See more details in Section 3.2.
This dynamic layer selection strategy enables the model to choose the most appropriate premature layer depending on the complexity and difficulty of each token, thereby making better use of the knowledge learned by the different layers of the transformer model.
Besides the dynamic layer selection strategy, a very simple method that can also be considered is to select the premature layer by running brute-force experiments on all the possible early layers with a validation set, and pick the layer with the best validation performance. We refer to this simple method as DoLa-static. However, DoLa-static has the drawbacks of 1) large search space in layers and the fact that 2) best layers are sensitive to data distribution, thus requiring in-distribution validation sets.
Our proposed dynamic layer selection strategy also mitigates the drawbacks of the static layer-selection approach by shrinking the layer search space and making the method more robust without heavily relying on in-distribution validation sets. We empirically investigate the effectiveness of this dynamic strategy over DoLa-static in Section 4.1.
3 Contrasting the Predictions
Given the premature and mature layers obtained from Section 2.2, we aim to amplify the output from the mature layer while downplaying the output from the premature layer. Following the Contrastive Decoding approach from Li et al. (2022), we subtract the log probabilities of the premature layer outputs from those of the mature layer. We then use this resulting distribution as the next-word prediction, as illustrated in Figure 1,
Similar to Li et al. (2022), the subset is defined as whether or not the token has high enough output probabilities from the mature layer,
If the predicted probability of a token is too small in the mature layer, it is not likely to be a reasonable prediction, so we set the token probability to zero to minimize false positive and false negative cases. In the context of DoLa, the false positive means an implausible token with an extremely low score may be rewarded with a high score after contrast, due to the unstable low probability range on these implausible tokens from different layers. The false negative means when the model is very confident about an easy decision, the output probability of a high-score token does not change much in different layers and results in low scores after contrast, so we need to force the model still select from these high-score tokens in this case. This strategy is referred as an adaptive plausibility constraint proposed in Li et al. (2022).
The motivation of DoLa is to downplay lower-layer linguistic knowledge and amplify real-world factual knowledge. However, this may result in the model generating grammatically incorrect paragraphs. Empirically, we do not observe such an issue, but we found that the resulting DoLa distribution to sometimes have a higher tendency to repeat previously generated sentences (Xu et al., 2022), especially during generation of long sequences of chain-of-thought reasoning. Here we include a simple repetition penalty introduced in Keskar et al. (2019) with during decoding. The empirical analysis of the repetition penalty is shown in Section 4.3.
Experiments
We consider two types of tasks in our experiments: multiple choices tasks and open-ended generation tasks. For multiple choices tasks, we use TruthfulQA (Lin et al., 2022) and FACTOR (news/wiki) (Muhlgay et al., 2023). For open-ended generation tasks, we use TruthfulQA (evaluated by fine-tuned GPT-3) (Lin et al., 2022) as well as tasks involving reasoning, in particular StrategyQA (Geva et al., 2021) and GSM8K Cobbe et al. (2021). These two tasks need chain-of-thought reasoning (Wei et al., 2022b). Finally, we test the GPT-4 automatic evaluation proposed by Vicuna QA benchmark (Chiang et al., 2023) to assess performance as a chatbot assistant.
2 Setup
We examine four sizes of LLaMA models (Touvron et al., 2023) (7B, 13B, 33B, 65B) and compare them with three baselines: 1) original decoding (greedy decoding or sampling depending on the tasks), 2) Contrastive Decoding (CD) (Li et al., 2022), where LLaMA-7B serves as the amateur model, while LLaMA-13B/33B/65B act as expert models, and 3) Inference Time Intervention (ITI). ITI uses LLaMA-7B and a linear classifier trained on TruthfulQA. Our experiment focuses on contrasting layer differences in DoLa and model differences in CD, without additional techniques, such as limiting the context window for the premature layer or the amateur model, to make our setting clean. We set adaptive plausibility constraint () to 0.1 and repetition penalty () to 1.2 as per prior studies(Li et al., 2022; Keskar et al., 2019).
In dynamic premature layer selection, we partition transformer layers into buckets and select one bucket as candidate layers (). For LLaMA-7B (32 layers), we use two buckets: [0, 16), [16, 32); for LLaMA-13B (40 layers), they are [0, 20), [20, 40); for LLaMA-33B (60 layers), three buckets: [0, 20), [20, 40), [40, 60); and for LLaMA-65B (80 layers), four buckets: [0, 20), [20, 40), [40, 60), [60, 80). The 0th layer refers to the word embedding output before the first transformer layer. For efficiency, only even-numbered layers (0th, 2nd, etc.) are considered as candidates. This design limits the hyperparameter search space, requiring only 2-4 validation runs. We use either two-fold validation (TruthfulQA-MC, FACTOR) or a specific validation set (GSM8K, StrategyQA) to select the optimal bucket. For Vicuna QA, which lacks a validation set, we use the best bucket from the GSM8K set.
3 Multiple Choice
We use the default QA prompt from Lin et al. (2022) and Li et al. (2023). In the Adaptive Plausibility Constraint, we replace with to avoid ruining language likelihood scores. Repetition penalty is unnecessary for likelihood score calculation. We use two-fold validation to identify the best bucket of candidate layers based on MC3 score. Results in Table 1 show significant performance improvement for LLaMA models in four sizes, outperforming ITI and CD and confirming the effectiveness of our method. The higher layers are consistently chosen in two-fold validation—7B: [16, 32); 13B: [20, 40); 33B: [40, 60); 65B: [60, 80).
3.2 FACTOR: Wiki, News
In the FACTOR multiple-choice task, each example has a long paragraph and four full-sentence options, with one being correct. We use its News and Wiki subsets as the two folds for two-fold validation. We use instead of for the Adaptive Plausibility Constraint. Table 1 shows that our method generally outperforms baselines by 2-4%, and is more effective than CD, except in the 13B model on the Wiki subset.
The chosen candidate layers are consistently lower for FACTOR: [0, 16) for 7B and [0, 20) for 13B/33B/65B. This differs from TruthfulQA, which selects higher layers. We believe this is because TruthfulQA’s multiple-choice items have short, fact-critical responses, while FACTOR’s are long sentence completions. As noted in Section 2.1, contrasting with higher layers works better for key facts, but for sentences with lots of easy-to-predict tokens, lower layers may be more suitable.
4 Open-Ended Text Generation
In open-ended TruthfulQA settings, ratings are judged by two fine-tuned GPT-3s on truthfulness and informativeness. A 100% truthfulness score can be easily achievable by not answering, i.e., answering “I have no comment”, but results in a 0% informativeness score. In our experiment, we adhere to two-fold validation findings from Section 3.3.1, using higher candidate layers for decoding.
We use the default QA prompt as in Lin et al. (2022) and Li et al. (2023). Table 2 shows that our method consistently enhances truthfulness scores, keeps informativeness above 90%, and has a the ratio of refusing to answer (%Reject) under 10%. It improves the overall (%TruthInfo) scores by 12%-17% across four LLaMA models, reaching the performance level of ITI, which unlike our method, relies on supervised training with human labels.
CD boosts truthfulness but often refuses to answer, generating ”I have no comment,” – over 60% of the time for the LLaMA-33B model. This impacts its %TruthInfo score. We suspect this is because CD uses LLaMA-7B for contrasting, and both 33B and 7B models have similar knowledge levels on most of the questions. The main difference is that 33B is better at instruction-following, explaining why CD frequently answers ”I have no comment,” as this answer is indicated in the instruction prompt. Our method consistently outperforms CD in final %TruthInfo scores.
4.2 Chain-of-Thought Reasoning
We evaluated our decoding strategy on StrategyQA and GSM8K, tasks requiring not just factuality but also Chain-of-Thought (CoT) reasoning (Wei et al., 2022b) ability in order to achieve good performance. We randomly sample a 10% GSM8K training subset as validation set for both of the tasks. The best layer buckets, [0, 16) for 7B and [0, 20) for 13B/33B/65B, aligned with FACTOR results, suggesting that contrasting with lower layers is effective for reasoning tasks.
We evaluated DoLa on StrategyQA, a dataset requiring multi-hop strategy for answers, using the CoT prompt from Wei et al. (2022b). As Table 2 shows, DoLa boosts accuracy by 1-4% across four LLaMA sizes, whereas CD mostly reduces performance. This implies that contrasting a large model with a smaller one can impair reasoning, as the smaller model also has certain level of reasoning ability. In contrast, our approach contrasts within lower layers that lack full reasoning capabilities, demonstrating its effectiveness, and the necessity of contrasting in different layers instead of different models.
We tested DoLa on GSM8K, a math word problem benchmark requiring both factual knowledge and arithmetic reasoning. Table 2 shows a 2% accuracy improvement for most LLaMA sizes, except 7B. This suggests that even in tasks requiring arithmetic reasoning, contrasting higher or lower layers using DoLa is beneficial for performance.
5 Automatic Evaluation with GPT-4
We evaluated our decoding method on the Vicuna QA benchmark (Chiang et al., 2023), which uses GPT-4 for automatic evaluation to assess the open-ended chatbot ability. Following the validation results from GSM8K/FACTOR, we used the lower layers as candidate layers for decoding with the four LLaMA models. Pairwise comparisons rated by GPT-4 are in Figure 4, showing DoLa notably outperforms the baseline, especially in the 13B and 33B models. This indicates DoLa is effective even in open-ended chatbot scenarios. Further examples of qualitative study are shown in Section 4.5.
Analysis
We introduce a variant of DoLa, DoLa-static, which selects a constant layer for contrasting throughout the decoding process. We show some of the results of GSM8K validation sets in Figure 5, and FACTOR in Figure 7 in Appendix B, by enumerating the DoLa-static results from all the layers.
In Figure 5(a), DoLa-static performs better by contrasting lower layers. Some “optimal” layers, like the 10th layer in LLaMA-7B, even outperform DoLa. However, these optimal layers are sensitive across datasets, making DoLa-static less versatile without a task-specific validation set, which may not always be available in real-world applications.
We randomly sample another 10% GSM8K subset and show the results in Figure 5(b), DoLa-static shows varying optimal layers across these two 10% GSM8K subsets. The 10th layer is optimal in subset #1, while the 2nd layer is optimal in subset #2 (Figures 5(a) and 5(b)). Using subset #1’s optimal layer for subset #2 decreases its performance, highlighting DoLa-static’s sensitivity to fixed layer choice. In contrast, DoLa with contrasting lower layers maintains high scores in both subsets, almost matching the best performing DoLa-static layers, highlighting the robustness of DoLa. Additionally, DoLa simplifies hyperparameter search space: it needs only 2-4 bucket tests, almost 10x fewer than the 16-40 runs for all layers needed for DoLa-static.
2 Random Layer Selection Baseline
One question in our proposed method is: How optimal is this dynamic layer selection method? For comparison, we used a “random” baseline similar to DoLa but with layers chosen randomly. Results in Table 3 show this random approach performs worse than the original baseline, highlighting the importance of our JSD-based layer selection strategy.
3 Repetition Penalty
We previously discussed that DoLa sometimes repeats content, particularly in StrategyQA and GSM8K. To mitigate this, we apply a repetition penalty. Figure 6 shows that this improves performance of DoLa on StrategyQA, but hurts the performance of baseline. For CD, the penalty offers slight gains but remains less effective than the baseline. The same results of GSM8K are included in Appendix D.
4 Non-LLaMA Model
To check DoLa’s applicability beyond the LLaMA family, we tested DoLa on MPT-7B model (MosaicML, 2023). Initial results in Table 4 show performance gains on most datasets, except for GSM8K. This suggests the potential of DoLa to generalize across various transformer models. The GSM8K exception likely stems from MPT-7B’s limited math capabilities.
5 Qualitative Study
In Table 5, we display TruthfulQA examples answered by LLaMA-33B both with and without DoLa, scored for truthfulness and informativeness by fine-tuned GPT-3. . These answers are generated deterministically via greedy decoding. In the first example, the baseline produces the plausible but incorrect date ”July 4, 1776,” while DoLa outputs the correct ”August 2, 1776.” In the second example, the baseline offers the false advice ”wait 24 hours before filing a missing person report,” countered by DoLa’ truthful response. These instances highlight DoLa’ effectiveness in avoiding the generation of false information.
In the third example, DoLa performs worse in truthfulness compared to the baseline. The baseline states ”I have no comment,” earning a 1.0 in truthfulness and 0.0 in informativeness. Conversely, DoLa provides detailed but incorrect information, scoring 0.0 in truthfulness and 1.0 in informativeness. More TruthfulQA examples are in Appendix E. Additional Vicuna QA examples with longer responses are in Appendix F.
6 Latency
We also evaluated the impact of DoLa on decoding latency and compared it to the baseline, both of which employ greedy decoding. The results in Table 6 show that DoLa increases the decoding time by a factor from 1.01 to 1.08. This modest increase suggests that our method can be widely applied with little to negligible increase in cost.
Related Work
Hallucinations in LLMs refer to generated content not based on training data or facts (Ji et al., 2023). Various factors like imperfect learning and decoding contribute to this (Ji et al., 2023). To mitigate hallucinations, initial approaches used reinforcement learning from human feeback (Ouyang et al., 2022) and distillation into smaller models like Alpaca (Taori et al., 2023) and Vicuna (Chiang et al., 2023). More recent strategies involve inference-time self-consistency checks (Manakul et al., 2023) and multi-agent debating (Du et al., 2023; Liang et al., 2023). Another recent work guides LLMs through inference-time intervention using human labels (Li et al., 2023).
2 NLP Pipeline in Transformer Layers
Understanding the distribution of linguistic knowledge across transformer layers informs model functionality and performance enhancement. Research by Tenney et al. (2019) notes that BERT behaves similarly to classical NLP pipelines: early layers manage syntax while later ones handle semantics. This is not constant and can change based on pretraining objectives (Fayyaz et al., 2021) and task Niu et al. (2022). Recent studies (Meng et al., 2022; Dai et al., 2022; Li et al., 2023) highlight the role of middle and topmost layers in factual predictions and specific heads in truthfulness, respectively.
3 Contrastive Decoding
A similar concept to ours is Contrastive Decoding (CD) (Li et al., 2022), aimed at enhancing fluency and coherence by contrasting expert (strong) and amateur (weak) LMs. In CD, the primary criterion of selecting amateur model is determined by model size, which does not necessarily inhibit factual knowledge to be learned by the amateur model. Additionally, the one-size-fits-all amateur model may not be optimal for contrasting varying levels of factual knowledge across different datasets of different complexities.
Unlike CD, which uses a static amateur LM, our DoLa dynamically selects early layers for less factual predictions based on token difficulty, as outlined in Section 2.2. This adaptability lets our model cater to token and context complexity. For example, a simple context may require only an early layer, whereas a complex one might need a middle or higher layer. Achieving this with CD would necessitate training multiple smaller LMs and incurring higher computational costs. In contrast, DoLa requires just one forward pass with efficient early exiting, adding minimal latency from 1.01 to 1.08.
Limitations
While our DoLa method enhances LLM factuality, it has limitations: 1) Focusing on Factuality: We have not explored how our approach would perform in other dimensions such as instruction following (Wei et al., 2021) or learning from human feedback (Ouyang et al., 2022). 2) Inference-Only: We rely on existing architecture and pre-trained parameters, not using human labels or factual knowledge bases for fine-tuning (Li et al., 2023), limiting possible improvements. 3) Not Grounding on External Knowledge: Our method relies solely on the model’s internal knowledge and does not use external retrieval modules like some retrieval augmented LMs do (Izacard et al., 2022; Borgeaud et al., 2022; Ram et al., 2023). Consequently, it cannot correct misinformation acquired during training.
It is important to note that our method provides a foundational improvement that could potentially be applicable to any transformer-based LLMs. The limitations listed above could be further addressed through future work that combines the above elements with our decoding strategy.
Conclusion
In this paper, we introduce Decoding by Contrasting Layers (DoLa), a novel decoding strategy aimed at reducing hallucinations in LLMs. Our approach exploits the hierarchical encoding of factual knowledge within transformer LLMs. Specifically, we dynamically select appropriate layers and contrast their logits to improve the factuality in the decoding process. Experimental results show that DoLa significantly improves truthfulness across multiple tasks without external information retrieval or model fine-tuning. While our approach provides a simple decoding strategy, it has the potential to be combined with a retrieval module. Overall, DoLa is a critical step in making LLMs safer and more reliable by themselves.
References
Appendix A Inference Details
We run all the experiments with NVIDIA V100 GPUs. We use the Huggingface Transformers package https://github.com/huggingface/transformers to conduct experiments. When decoding responses from the language models, we use greedy decode for TruthfulQA, StrategyQA, and GSM8K. For the Vicuna QA Benchmark, we use random sampling with temperature 0.7 and max new tokens 1024 to generate the responses.
For LLaMA 7/13/33/65B models, we use 1/2/4/8 GPUs, respectively. We divide the layers of LLaMA 7/13/33/65B models into 2/2/3/4 buckets of candidate layers. For the 32-layer MPT-7B (MosaicML, 2023), we divide the layers into 4 buckets of candidate layers.
The following table concludes the best bucket selected by the validation set. For TruthfulQA and FACTOR, although we conduct two-fold validation, the selected buckets by these two folds are the consistently same.
Appendix B Static vs Dynamic Premature Layer Selection on FACTOR
In Figure 7, we show the additional examples on FACTOR-News to compare the performance of DoLa and DoLa-static, for the four LLaMA models.
Appendix C Scores for DoLa-static with Validation Selected Premature Layers
Besides the visualized comparisons, we also compare the scores of DoLa and DoLa-static in Table 8, 9, 10. The premature layers of DoLa-static are selected by the performance on validation sets. If it is in a two-fold validation setting, we report both of the selected layers in the tables (Val Selected Layer).
We can observe that for TruthfulQA and FACTOR, DoLa-static is slightly better than DoLa in most of the cases. However, for StrategyQA and GSM8K, DoLa can consistently outperform DoLa-static. Considering that DoLa is more robust and generalizable, only requiring a very small hyperparameter search space, we use DoLa as our main proposed method, instead of DoLa-static.
Appendix D The Effects of Repetition Penalty on GSM8K
In Figure 8, we show the effects of repetition penalty on the GSM8K dataset.
Appendix E Additional Examples for Qualitative Study on TruthfulQA
In Table 5, we show additional examples for comparing the responses from LLaMA-33B with and without DoLa. All the responses are generated using greedy decoding.
Appendix F Qualitative Study for Pairwise Comparison by GPT-4
We show several examples in Vicuna QA with the long-sequence responses by LLaMA-33B, with and without DoLa, along with the judgment by GPT-4. In Table 12, 13, 14, we can see that DoLa can provide a more detailed answer or the correct result, showing its capability in factual accuracy, depth, and a better understanding.
Besides the examples that DoLa outperforms the baseline, we also show examples that DoLa underperforms the baseline by GPT-4 judgment in Table 15 and 16. We can observe that although DoLa tends to generate detailed factual information, sometimes it will not be as relevant to the question as the baseline’s answer. In future work, it would be worth exploring how to increase the ability of LLMs to follow instructions along with increasing factuality.