LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning

Hongye Jin, Xiaotian Han, Jingfeng Yang, Zhimeng Jiang, Zirui Liu, Chia-Yuan Chang, Huiyuan Chen, Xia Hu

Introduction

The context window length of most existing LLMs (Zhao et al., 2023; Yang et al., 2023) is limited since they are trained with a fixed length of training sequences. It’s determined by the context window length during the pretraining stage. Once the length of the input texts exceeds the pretraining context window during the inference, the behavior of LLMs will be unpredictable and suffer from severe performance degradation. The perplexity (PPL) of the model will explode with the long input sequences (Xiao et al., 2023; Peng et al., 2023; Han et al., 2023; Chen et al., 2023b).

Recently, a variety of context window extension methods have been developed to tackle the challenge of extending the context window size of pretrained LLMs. A straightforward approach is to fine-tune these models on enough extensive texts. Besides this, some methods seek to extend context window length in more efficient fine-tuning ways. Among these contemporary methods, some notable techniques include ‘PI’ (Chen et al., 2023b), ‘CLEX’ (Chen et al., 2023a) ‘Yarn’ (Peng et al., 2023), ‘LongLora’ (Chen et al., 2023c), and ‘ABF’ (Xiong et al., 2023). These methods aim to extend the content window based on the implicit assumption that pretrained LLMs lack the ability to handle long content. However, these methods typically require finetuning to achieve extension, which can be resource and time intensive given the quadratic complexity of Transformers. Additionally, high-quality long text data is scarce, hindering such fine-tuning approaches. Most real-world data is short, and much long text lacks meaningful long-range dependencies. With limited appropriate data, finetuning risks degrading existing strong performance on shorter sequences from pretraining or overfitting models to the tuning set. LLMs’ generalizability to broad tasks may reduce.

Instead of extending the content window, in this paper, we believe LLMs should have inherent capabilities to handle long contexts. Our belief stems from the fact that when we, as human beings, are children, we are taught how to read and write using relatively short texts, such as articles spanning several pages. We rarely use extremely long texts like entire books or complete documents as learning materials. Yet, we are still able to understand long texts effectively. With this strong motivation, the poor performance of LLMs while facing long text out of the pretraining context window is not due to the lack of long context understanding capabilities.

In our analysis, the key challenge preventing LLMs from effectively handling longer contexts is the Out-of-Distribution (O.O.D) issues related to positional encoding, which we call the positional O.O.DHere, the position refers to relative position rather than absolute position. The relative position is m−nm-n in RoPE, where mm and nn are the absolute positions of two tokens. The positional O.O.D refers to cases where the value of m−nm-n during inference is unseen, i.e., larger than the values observed during pretraining. In this paper, we map unseen large relative positions to those observed during pretraining. More details about m−nm-n are provided in Section 2. issue. This problem arises when LLMs encounter text sequences during inference exceeding the length of their pretraining context window, where LLMs are exposed to new relative distances that were not present during their pretraining phase. It is widely recognized that Neural Networks (NNs) are susceptible to unpredictable behaviors when dealing with O.O.D inputs (Liu et al., 2021; Shen et al., 2021; Bai et al., 2021; Zhang et al., 2023). To address this, an intuitive and practical solution would be to remap the unseen relative positions to those encountered during the pretraining, thus extending the LLMs’ ability to handle longer contexts naturally.

This paper proposes SelfExtend to elicit LLMs’ inherent long context capabilities. SelfExtend addresses the issue of O.O.D. positional information by using a simple floor division operation to map unseen large relative positions to those encountered during pretraining. The core idea hinges on the observation that, in long texts, exacting word positions becomes less crucial. The overall meaning and the relative order of information hold greater significance. Just like when answering questions about lengthy texts, we rely on the general location and order, not the specific word-by-word placement. Natural language exhibits a characteristic where meaning stays relatively consistent within short ranges like paragraphs. Therefore, using close or even identical position encodings effectively captures the necessary relative ordering of important information. This intuitive approach aligns perfectly with the floor operation’s functionality. Additionally, T5 (Raffel et al., 2020) and iRPE (Wu et al., 2021) also share this similar intuition.

Our SelfExtend is a plug-and-play method that takes effect at the inference stage, allowing existing large language models to easily adopt it. We evaluate SelfExtend with some popular LLMs (Llama-2 (Touvron et al., 2023), Mistral (Jiang et al., 2023), SOLAR (Kim et al., 2023), and Phi-2 (Javaheripi et al., 2023)) on three types of tasks: language modeling, synthetic long context tasks, and real-world long context tasks. The proposed SelfExtend substantially improves the long context understanding ability and even outperforms many finetuning-based methods on some tasks. These results underscore SelfExtend as an effective solution for context window extension. The superior performance of SelfExtend also demonstrated the potential of large language models to effectively handle long contexts. Our main contributions are summarized as follows:

We think LLMs with RoPE have a natural ability to handle long texts, even if they have not encountered super-long ones during training. The previous limitation stems from O.O.D. positions, meaning the ”larger” positions have not been seen during training. We call this the positional O.O.D. issue.

Based on this belief and to address the positional O.O.D. issue, we propose SelfExtend to extend the context window of LLMs without any fine-tuning. We map the unseen large relative positions (at inference) to known positions (at training), thus allowing LLMs to maintain coherence over longer texts without additional fine-tuning.

In both synthetic and real-world long context tasks, SelfExtend has proven its ability to deliver performance that matches or surprisingly surpasses many existing fine-tuning-based models. This highlights the superior capabilities of our SelfExtend model.

Preliminary

Position Encoding. Transformers (Vaswani et al., 2017) incorporate position information via different positional embedding designs. The positional embedding design can be categorized into two classes: absolute position embeddings and relative positional encodings. The absolute position embedding provides the absolute positions, which embeds each absolute position ii into position vector pi\mathbf{p}_{i} and adds word embeddings to their corresponding pi\mathbf{p}_{i} before feeding them to the model. Examples of such include sinusoidal position embeddings (Vaswani et al., 2017) and learned position embeddings in GPT3 (Brown et al., 2020) and OPT (Zhang et al., 2022), or adding the dot product between two tokens’ position embeddings on the attention logit (Ke et al., 2020).

On the other hand, relative positional encodings have been proposed to use relative distance information between tokens and have become the mainstream of position embedding. This information is usually applied in attention layers. Examples of such include a learnable attention logit bias as in T5 (Xue et al., 2020), Transformer-XL (Dai et al., 2019); a fixed linear attention decay called Alibi (Press et al., 2021); rotating query and key sequences based on distance such as RoPE (Su et al., 2022), and XPos (Sun et al., 2023). The proposed method is based on the Rotary Position Embedding (RoPE) introduced in (Su et al., 2022).

where g(⋅)g(\cdot) is an abstract mapping function.

SelfExtend

In this section, we first conduct a preliminary investigation on the inherent ability of the LLMs to handle long content. Then, we propose our SelfExtend that effectively extends existing LLMs’ context window without any fine-tuning.

① Why do LLMs fail on sequences during inference that are longer than their pre-training context window? For a pretrained LLM with relative position encodings, such as RoPE, the behavior of the LLMs becomes unpredictable during inference if the length of a sequence is longer than its pretraining context window length. This has been explored by (Han et al., 2023; Chen et al., 2023b) that with unseen relative positions, the attention distributions are very different compared to those within the pretraining context window length. We argue that such failure stems from the Out-of-Distribution (O.O.D.) relative distance in the same sense that neural networks are not robust to O.O.D. inputs (Shen et al., 2021).

② How to solve positional O.O.D. problem? One feasible and straightforward way to handle unseen relative positions is to map them to positions that were seen during pretraining. We can use the floor operation to map the unseen positions to positions within the pretraining context window, as shown in Figure 1. The proposed method is identical to the original self-attention mechanism except that the floor operation is applied to each token’s original position before the inner product. We denote the self attention with the floor operation applied as “grouped attention”. In Python style, the “grouped attention” is denoted as:

③ Can LLMs work well without accurate position information? — Yes, but not that perfect. We show the perplexity (PPL) on the PG-19 (Rae et al., 2019) dataset with the floor operation applied to Llama-2-7b-chat across different sequence lengths, in Figure 2. As a comparison, we also show the PPL of the original model without the floor operation as the dotted lines. From this figure, with the floor operation, LLMs keep a relatively low and stable PPL on the sequences whose lengths exceed the pretraining context window. Meanwhile, with grouped attention, the PPL is a little higher than the original LLMs, which is expected. However, the model’s PPL behavior is similar to the original model, as the PPL is nearly unchanged within the “context window” (for Llama-2: 2 - 8192, 4 - 16384, and 8 - 32768), demonstrating the effectiveness of group attention. Once the length of a sequence is longer than the new “context window” (e.g., sequences with 10k tokens as the input, with a group size of 2 ), the PPL explodes again due to the positional O.O.D issue.

④ How to restore degraded language modeling ability caused by grouped attention? — Re-introducing normal attention in the neighboring area. In the process of generating next tokens, the immediate neighbors of a target token play a crucial role, which is well-supported by existing methods of sparse attention mechanisms (Zaheer et al., 2020; Shi et al., 2021) and methods for extending the contextual window (Han et al., 2023; Xiong et al., 2023; Chen et al., 2023c). These studies consistently highlight the importance of maintaining the standard attention mechanism for tokens in close proximity to the target token. This proximity-based focus is essential for the accurate generation of the next token, ensuring the coherence and fluency of the generated text, as evidenced by acceptable perplexity (PPL) levels. Employing grouped attention might not significantly affect the overall quality of generated sentences; however, it necessitates the accurate positioning of attention to maintain generation quality. Therefore, it is imperative to preserve the standard attention mechanism within the vicinity of the target token, as utilized during the pretraining phase, to ensure the precision and effectiveness of language models in capturing the nuances of local context.

2 SelfExtend LLM Context Window Without Tuning

We introduce SelfExtend, a method that enhances LLMs’ natural capability to process extensive contexts without the need for fine-tuning. SelfExtend incorporates two distinct types of attention mechanisms: 1) Grouped attention, specifically designed for tokens that are far apart. This approach applies a floor operation to the positions to manage long-distance relationships between tokens; 2) Standard attention, which employs the conventional attention mechanism for adjacent tokens within a specified range. The SelfExtend framework is depicted in Figure 3. Notably, SelfExtend modifies only the attention mechanism during inference, eliminating the need for additional fine-tuning.

Maximum Extended Length of SelfExtend Suppose that we have the pretraining context window size as LL, the group size for grouped attention as GsG_{s}, and the window size for neighbor tokens as wnw_{n}. We shift the relative position of grouped attention by wn−wn//Gsw_{n}-w_{n}//G_{s} before merging the two pieces of attention together. This ensures that the transition from the normal attention area to the grouped attention area smooth. We merge the two parts of attention by replacing the attention values out of the neighbor token window with the attention values from the grouped attention. All the modifications are applied before the softmax operation and other parts remain unchanged. Ideally, the maximum length of the extended context window is:

For example, in Figure 3, the context window is extended from its pretraining length of 77 to (7−4)∗2+4=10(7-4)*2+4=10. The pseudo code for SelfExtend are presented in Algorithm 1.

Relation to Existing Work The grouped attention in SelfExtend can be viewed as a form of position interpolation (Chen et al., 2023b), where some positions are interpolated to be infinitely close to pretraining positions. Another finetuning-free method, ReRoPE (Su, 2023), is equivalent to a special case of SelfExtend: the group size is large enough that all tokens outside the neighbor window fall into the same group (e.g. group size 10,000 in Figure 5). T5 (Raffel et al., 2020) and iRPE (Wu et al., 2021) also share the high-level idea of multi-level positional encodings, while applying it during pretraining. T5 is more similar to ReRoPE for using the same position for distant tokens. iRPE has finer distant position encodings, more akin to SelfExtend.

Experiments

We evaluate SelfExtend with Llama-2 (Touvron et al., 2023) and its families, Phi-2 (Javaheripi et al., 2023), Mistral (Jiang et al., 2023) and SOLAR (Kim et al., 2023) on language modeling task, synthetic long context tasks, real-world long context tasks and standard short-context tasks.

Language modeling task is the most fundamental and the least requirement for LLMs, which is usually measured by perplexity (PPL) on the test text data. A low PPL does not guarantee good performance on real tasks (Pal et al., 2023), however, a higher PPL suggests severe performance degradation of LLMs. We evaluate SelfExtend’s language modeling performance on dataset PG19 (Rae et al., 2019), which contains lengthy books. PPL is used as the metric. More experimental details are presented in Section D.1

The results show that SelfExtend can successfully maintain a low PPL out of the pretraining context window for both Llama-2-7b-chat and Mistral. Without SelfExtend, the PPL explodes when the length of test sequence is larger than the context window. Mistral with SWA can also maintain a low PPL out of its context window. But later in the next section, we will demonstrate that a low PPL score does not necessarily indicate proficiency in handling long contexts. More discussion about PPL can be found in Appendix B.

2 Performance on Synthetic Long Context Tasks

The passkey retrieval task is the same as what is defined in Landmark Attention (Mohtashami & Jaggi, 2023), which is a synthetic long context task. It requires a language model to retrieve a simple passkey (i.e., a 5-digit random number) in a long meaningless text sequence. The passkey is placed with various document depths (where the passkey is placed in the input texts) and context lengths (ranging from 4k to 24k). We tested multiple passkey retrievals for each context length and depth. The passkey was randomly placed within a span of 400400 tokens. For a depth of 0.10.1 and context of 8k, the passkey was placed between tokens 800−1600800-1600. We performed 1010 iterations per span, so 2020 total for that setting. Experimental setting details and an example of passkey retrieval task can be found in Section D.2.

The results in Figure 4 show that without any fine-tuning, SelfExtend obtains 100% passkey retrieval accuracy across all tested depths and context lengths. The results also demonstrate that: although Mistral w/ SWA has low PPL beyond its pretraining context window, it can only access information (i.e. the passkey) within its sliding window. Considering the simplicity of this task, these results strongly suggest it still does not have the true ability to handle long contexts.

3 Performance on Real-World Long Context Tasks

Evaluation solely on language modeling (measured by perplexity) and synthetic tasks like passkey retrieval cannot fully assess the long-context capabilities of LLMs. The task of Passkey retrieval is overly straightforward, and an LLM may still struggle with long context despite low perplexity. To comprehensively evaluate long-context performance, we further use two recent real-world long context benchmarks: LongBench (Bai et al., 2023) and L-Eval (An et al., 2023). The results are presented in Table 2 and Table 3.

On the LongBench in Table 2, for all four different base LLMs and most datasets, with SelfExtend, the LLM can obtain significant performance improvments. Llama-2-7B: We use SelfExtend to increase Llama-2-7b-chat’s context from 4k to 16k and 25k. Both significantly outperform Llama-2-7b-chat and most fine-tuned models on several datasets like HotpotQA. We also extend vicuna1.5-7B from 4k to 16k and 25k. With SelfExtend, vicuna1.5-7B surpasses its fine-tuned counterpart vicuna1.5-7B-16k and ranks among top Llama-2-7b models. On some datasets, the 25k variant underperforms the 16k one due to the trade-off between larger context and positional precision. More details about the trade-off is in Section 4.5. Mistral-7B: We extend Mistral-7B’s context to 16k, significantly improving its long context ability over the base model, with or without SWA applied. The fine-tuned variant MistralLite ((amazon, 2023)) achieves the best performance on most datasets. However, many of these datasets were included in MistralLite’s fine-tuning data, such as NarrativeQAMore details about MistralLite’s fine-tuning data can be found at https://huggingface.co/amazon/MistralLite. At least, GovReport, QMSum, NarrativeQA, Qasper, QuALITY, and HotpotQA are included. Meanwhile, Multi-passage QA and summarization tasks are also in fine-tuning data. This also violates zero-shot evaluation conditions.. SOLAR-10.7B and Phi-2: They have no finetuned variant for context window extension yet. SelfExtend can also obtain substantial performance improvements.

On the LEval benchmark in Table 3, we observe similar results. Compared to fine-tuning free baselines like NTK or further fine-tuned models like Longchat1.5-7b-32k and Vicuna1.5-7b-32k, SelfExtend achieves superior performance on nearly all datasetsLEval performance seems sensitive to prompt engineering for these sub-13B LLMs. For example, on some datasets, vanilla vicuna-13b underperforms vanilla vicuna-7b..

In summary, on the two benchmarks, SelfExtend achieves comparable or better performance, compared to methods that requires further fine-tuning. Despite our initial expectation being that SelfExtend would simply outperform the base model without additional extension methods, it is remarkable that our SelfExtend, which solely operates during inference without the need for fine-tuning or training, achieves such impressive performance.

4 Performance on Short Context Tasks

We argue that an ideal context length extension method should not degrade performance on standard short-context tasks. Previous fine-tuning based methods usually undergo performance degradation on short-context tasks (Peng et al., 2023; Xiong et al., 2023). Following (Peng et al., 2023), we use Hugging Face Open LLM Leaderboard (Gao et al., 2023) to evaluate SelfExtend’s performance on five public short context tasks. Specifically, we use 25-shot ARC-Challenge (Clark et al., 2018), 10-shot HellaSwag (Zellers et al., 2019), 5-shot MMLU (Hendrycks et al., 2020), 0-shot TruthfulQA (Lin et al., 2021), and 5-shot GSM8K (Cobbe et al., 2021). The results are shown in Table 4. We also investigate the influence of varying group sizes and neighbor window sizes on short-context tasks and we present the results in Appendix C.

The results show that SelfExtend can maintain the performance of the short-context tasks, while enhance the performance on long-context tasks. Moreover, because SeldExtend does not require any fine-tuning and only takes effect during inference, SelfExtend can be readily adopted as a plug-in component for LLMs. This means SelfExtend can be automatically and inherently disabled while encountering short-text sequences. Then, with the parameters remaining unchanged, LLMs can maintain its original inference mechanism on those short-context scenarios.

5 Ablations on Group Size and Neighbor Window

We investigate the influence of varying the group size GsG_{s} and the neighbor window wnw_{n}. We experiments with Phi-2 on four real-world datasets from Longbench: narrativeqa, qasper, triviaqa, and repobench-p. The results are presented in Figure 5. Form the results, we observe two trade-offs:

1) There is a trade-off with respect to group size in SelfExtend. Generally, both too small and too large group sizes can result in inferior performance compared to an optimal level. With a large group size, position information becomes more coarse, potentially causing performance drops. Conversely, small group sizes require SelfExtend to utilize larger position embeddings to extend the context window. These larger position embeddings are less trained compared to smaller ones. For example, in Llama-2 with its 4096 context window, the relative position 4095 accounts for only 1/2048 the frequency of the relative position 2048 in training. These under-trained relative positions can also degrade performance. This trade-off produces the ’peak’ shape in the figure, indicating the extended context window differs from the ideal case described in Equation 4.

2) There is also another trade-off w.r.t. neighbor window size. With larger neighbor window sizes, there is more precise information about neighbor tokens, which is the most important. But a larger neighbor window size means SelfExtend has to use a larger group size for a long sequence, compared to using a smaller neighbor window size & smaller group size, the information about the whole sequence becomes coarse.

6 Performance with Varying Context Window Length

To validate SelfExtend’s efficacy in enabling LLMs to utilize extended context windows, we assess Phi-2’s performance across varying context lengths with SelfExtend, referencing Table 5. Across four task types from LongBench, results are generally improved with longer contexts. Notably, SelfExtend monotonically enhances performance on NarrativeQA and Qmsum. While significant improvements are observed across most datasets, a ’peak’ in performance suggests a trade-off, as discussed in Section 4.5: longer contexts offer more relevant information, but the larger group sizes required by SelfExtend to extend the context window may cause less precise positional informationOther possible reasons include: Phi-2 is a base model without instruction tuning, and SelfExtend’s performance is not optimal as we use the same set of hyperparameters across all datasets, which cannot showcase SelfExtend’s full potential. Regarding Lcc, performance remains consistent, possibly due to its reliance on local codes and shorter dataset lengthsWith Phi-2 tokenizer, over 60%60\% of Lcc instances are under 4096 tokens, with an average length of 4069.7.

7 Varying-Length Passkey Retrieval Task

The conventional passkey retrieval task, along with prevalent benchmark datasets, primarily assesses the proficiency of LLMs in identifying and leveraging pertinent information. Traditionally, this task involves passkeys not exceeding 5 digits in length. To evaluate the LLMs’ capabilities of producing consistent and precise outcomes for long sequences, we extended the task to incorporate passkeys with larger lengths. We test passkeys in 5,8,16,36,48,64,1005,8,16,36,48,64,100 digits. The input sequence contains 16,00016,000 characters. More details are presented in Section D.3.

The results, depicted in Figure 6, illustrate a common trend: while short passkeys of 5 or 8 digits are easily managed by all, divergences in performance emerge as the length of passkey increases. Notably, with the exception of Yarn, many tuning-based methods are unable to accurately reproduce passkeys beyond 64 digits, and some of them even experience a marked decline in performance when the passkey length exceeds 16 digits. Remarkably, although without tuning, SelfExtend maintains its superiority. These findings suggest that we should carefully choose the training approach when fine-tuning models to handle long contexts.

Conclusion and Discussion

In this paper, we argue that LLMs themselves have the inherent ability to handle long sequences and propose SelfExtend to elicit the inherent long context abilities for LLMs by simply mapping unseen relative positions into those seen during pretraining via the Floor operation. Without any tuning or further training, SelfExtend can effectively improve LLMs’ long context performance, as extensive experiments show.

Limitations: SelfExtend increases computation cost with naive implementations since it performs extra attention across all query-key pairs. However, with optimizations like blocked kernels (e.g., Flash Attention (Dao et al., 2022)), this becomes linear rather than quadratic, and the marginal cost is small enough to be ignored for long input sequences. Also, the performance degrades with large group size, preventing indefinitely long contexts. Additionally, evaluation methodologies for assessing long context abilities remain open research questions.

Future Work: We are interested in testing SelfExtend on models using other positional encoding. Larger models, longer contexts, and more challenging tasks will be tested if we can access more computational resources in the future. In the meantime, more sophisticated mapping methods will be considered as the replacement of the simple floor operation to achieve better long context understanding abilities and extended context window length.

References

Appendix A Pseudocode of SelfExtend

Appendix B Perplexity as a Metric for Long Context Capabilities

PPL is not an effective metric for measuring the ability of LLMs to handle long contexts. In Figure 7, we introduce a seemingly plausible context window extension method named ’Infinite’. When evaluated on PG19 using the same protocol, Llama-2-7b-chat with ’Infinite’ achieves PPL scores that are comparable to, or even lower than, those achieved by SelfExtend, as demonstrated in Table 6. However, ’Infinite’ essentially mimics the process of dividing a long sequence into short sub-sequences before processing them with LLMs, indicating that it does not genuinely address long context handling.

The discrepancy between Perplexity (PPL) and long context ability primarily stems from how PPL is calculated by averaging over numerous tokens. As long as the majority of tokens are modeled accurately, PPL will remain low. This is closely related to the influence of neighboring tokens. Information from neighboring tokens—such as those within the local attention window of ’Infinite’—can suffice for predicting most tokens, thus leading to a low PPL. However, a few critical tokens, which are crucial for understanding long contexts and answering questions, may not be predicted accurately.

Additionally, unlike the pre-training process where the cross-entropy loss corresponds directly to perplexity, measuring PPL during inference is static. It resembles a specific point on the loss curve observed during pre-training. While a decreasing trend in loss during pre-training indicates good performance, a single point on the training loss curve cannot determine the performance.

In summary, while low PPL is essential for a good model, lower PPL does not necessarily equate to better performance in understanding long contexts.

Appendix C SelfExtend on Group Size and Neighbor Window

To comprehensively understand SelfExtend’s influence on LLMs, unlike previous experiments which used long context settings, we evaluate with smaller neighbor window sizes on four standard benchmark tasks: ARC-c, GSM8k, Hellaswag and MMLU. We use Phi-2 as the extended LLM. The results are shown in Figure 8. We didn’t include TruthfulQA because its average length is less than 300 words, while the four datasets we used have an average length greater than 700 words. In general, SelfExtend has a minor influence on Phi-2 as long as the neighbor window size is over 128. In many cases, SelfExtend even performs slightly better than vanilla Phi-2. When the neighbor window is too small (e.g. 64 tokens), if the group size is large, as expected, the positional information loss is too high and Phi-2’s performance degrades. Also, on difficult tasks such as MMLU and Helleswag, we observe a monotonic decrease in performance with increasing group size for all neighbor windows. In summary, even when applying SelfExtend to short context tasks, as long as the hyperparameters are not extreme, SelfExtend does not harm the model.

Appendix D Detailed Experimental Setting

In this appendix, we present the details of the experiments in our paper.

This is not the standard setting for PPL testing on PG-19. We use the first sentence of each book in PG19’s test set (100 books) to test the language modeling ability. The results cannot be directly compared to the PPL reported by other papers. We chose this setting because our computation resources are very limited. This setting saves a lot and it can still show the behavior of LLMs w.r.t. PPL. All PPL results were calculated using the sliding window method (Press et al., 2021) with S=256S=256. We evaluated how the PPL changes as the input length increases. In Table 1, SelfExtend extends the original Llama-2’s context window length from 40964096 (4k) to over 1638416384 (16k) with group size GsG_{s} set as 88 and neighbor window wnw_{n} set as 10241024 (1k). For Mistral model, without SWA, the context window is 81928192 (8k) and it is also extended by SelfExtend with the same setting to larger than 16k. With SWA, Mistral can digest an infinite length of sequences and its default sliding window is 40964096.

D.2 Experimental Setting on Passkey Retrieval Task

Compared to other synthetic tasks, such as ”Needle in a Haystack” (gkamradt, 2023), the model’s performance on this is not sensitive to the prompt (Anothropic, 2023). This may come from the fact that the sentence carrying the passkey is very different from those repeated random texts surrounding it. Empirically, within the effective context window, almost all LLMs, including those without any instruction tuning or alignment, can locate the sentence carrying the passkey. Although this task is easy and far from real-world scenarios, it tests two fundamental capabilities of LLMs: 1. The model should be able to recognize and locate the useful information across all positions of the input sequence (the most fundamental understanding capability); 2. The model should be able to use the perceived information to finish tasks (the most fundamental generation capability).

An example of passkey is as the following:

D.3 Experimental Setting on Varying-Length Passkey Retrieval Task

In this experiment, we use the following models: Llama2-7b-chat with SelfExtend, LongLora-7b-16kWe use its fully fine-tuned variant. For more details about the model: https://huggingface.co/Yukang/Llama-2-7b-longlora-16k. We didn’t use the latest version named ’LongAlpaca’ as we cannot get reasonable performance for this specific task with LongAlpaca, which may be due to our improper configuration.,vicuna-1.5-7b-16k, Together AI’s Llama-2-7b-32kBoth vicuna-1.5-7b-16k and Together AI’s Llama-2-7b-32k were fine-tuned using position interpolation, and Yarn-Llama-2-7b-64k.

Appendix E Details of LLMs

Here, we list the links to the details of the LLMs utilized in our experiments.