Effective Long-Context Scaling of Foundation Models

Wenhan Xiong, Jingyu Liu, Igor Molybog, Hejia Zhang, Prajjwal Bhargava, Rui Hou, Louis Martin, Rashi Rungta, Karthik Abinav Sankararaman, Barlas Oguz, Madian Khabsa, Han Fang, Yashar Mehdad, Sharan Narang, Kshitiz Malik, Angela Fan, Shruti Bhosale, Sergey Edunov, Mike Lewis, Sinong Wang, Hao Ma

Introduction

Large language models (LLMs), trained with an unprecedented magnitude of data and compute, hold the promise of fundamentally improving the way we interact with the digital world. As LLMs get rapidly deployed and continue to evolve through scaling, we envision these models to serve more intricate and complex use cases, such as analyzing dense knowledge-rich documents, powering more genuine and engaging chatbot experiences, and aiding human users in iterative creation processes such as coding and design, etc. A crucial feature supporting this evolution is the ability to effectively process long-context inputs.

Until now, LLMs with robust long-context capabilities are primarily provided through proprietary LLM APIs (Anthropic 2023; OpenAI 2023) and there is no open recipe for building long-context model that can demonstrate on-par downstream performance as these proprietary models. Moreover, existing open-sourced long-context models (Tworkowski et al. 2023b; Chen et al. 2023; Mohtashami and Jaggi 2023; MosaicML 2023b) often fall short on evaluations and primarily measure long-context capabilities with the language modeling loss and synthetic tasks, which do not comprehensively demonstrate their effectiveness in diverse, real-world scenarios. Additionally, these models often overlook the necessity of maintaining strong performance on standard short-context tasks, either bypassing the evaluations or reporting degenerated performance (Peng et al. 2023; Chen et al. 2023). † Equal contribution ∗ Corresponding authors:{xwhan, sinongwang, haom}@meta.com

In this work, we describe our approach to build long-context LLMs with superior performance over all existing open-sourced models. We build our models by continually pretraining from Llama 2 checkpoints with additional 400 billion tokens formed as long training sequences. Among the model series, the smaller 7B/13B variants are trained with 32,768-token sequences while the 34B/70B variants with 16,384-token sequences. In contrast to the limited evaluation performed by existing studies, we extensively evaluate our models using language modeling, synthetic tasks, and also a wide range of real-world benchmarks covering both long and short context tasks. On language modeling, our model demonstrates a clear power-law scaling behavior with respect to context lengths. This scaling behavior, as shown in Figure 1, not only shows our models’ ability to consistently benefit from more contexts but also suggest that context length is another importance axis of scaling LLMs. When comparing our models to Llama 2 on research benchmarks, we not only observe significant improvements on long-context tasks but also modest improvements on standard short-context tasks, especially on coding, math, and knowledge benchmarks. We explored using a simple and cost-effective procedure to instruction finetune our continually pretrained long models without any human-annotated data. The end result is a chat model that can achieve stronger overall performance than gpt-3.5-turbo-16k on a series of long-context benchmarks covering question answering, summarization, and multi-document aggregation tasks.

In the remaining part of this paper, we begin by presenting the continual long-context pretraining approach and a lightweight instruction tuning procedure, followed by detailed results on a range of short and long context tasks. To facilitate future research, we complement our results with an analysis section discussing how the design of positional encodings, the length distribution of the dataset and the training curriculum contributes to the final performance. Finally, we report responsible safety evaluations, which validates that our models can largely maintain the safety performance of the original Llama 2 series.

Method

Training with longer sequence lengths can introduce significant computational overhead due to the quadratic attention calculations. This is the main motivation of our continual pretraining approach. The underlying hypothesis that similar long-context capabilities can be learned by continually pretraining from a short-context model is later validated in Section 4.4 through comparing different training curricula. We keep the original Llama 2 architecture nearly intact for continual pretraining and only make a necessary modification to the positional encoding that is crucial for the model to attend longer. We also choose not to apply sparse attention (Child et al. 2019) in this work, since given Llama 2 70B’s model dimension (hh = 8192), the cost of attention matrix calculation and value aggregation only becomes a computation bottleneck when the sequence length exceeds 49,152 (6h6h) tokens (Narayanan et al. 2021). While sparse attention might be useful for reducing the key/value cache size at inference time when trading off performance, it can complicate the inference pipeline and the improvements can also be offset by quantization methods.

Through early experiments at the 7B scale, we identified a key limitation of Llama 2’s positional encoding (PE) that prevents the attention module from aggregating information of distant tokens. We adopt a minimal yet necessary modification on the RoPE positional encoding (Su et al. 2022) for long-context modeling – decreasing the rotation angle (controlled by the hyperparameter “base frequency bb”), which reduces the decaying effect of RoPE for distant tokens. In Section 4.1, we show this simple method outperforms a concurrent approach (Chen et al. 2023) for extending Llama’s context length and provide a theoretic explanation of its superiority.

Data Mix

On top of the working model with the modified PE, we further explored different pretrain data mixes in Section 4.2 for improving long-context abilities, either by adjusting the ratio of Llama 2’s pretraining data or adding new long text data. We found that often the quality of the data plays a more critical role than the length of texts for long-context continual pretraining.

Optimization Details

We continually pretrain Llama 2 checkpoints with increased sequence length while keeping the same number of tokens per batch as in Llama 2. We train all models for a total of 400B tokens over 100,000 steps. With FlashAttention (Dao et al. 2022), there is negligible GPU memory overhead as we increase the sequence length and we observe around 17%17\% speed loss when increasing the sequence length from 4,096 to 16,384 for the 70B model. For the 7B/13B models, we use learning rate 2e−52e^{-5} and a cosine learning rate schedule with 2000 warm-up steps. For the larger 34B/70B models, we find it important to set a smaller learning rate (1e−51e^{-5}) to get monotonically decreasing validation losses.

2 Instruction Tuning

Collecting human demonstration and preference labels for LLM alignment is a cumbersome and expensive process (Ouyang et al. 2022; Touvron et al. 2023). The challenge and cost are more pronounced under long-context scenarios, which often involve complex information flow and specialized knowledge, e.g., processing dense legal/scientific documents, making the annotation task nontrivial even for skilled annotators. In fact, most existing open-source instruction datasets (Conover et al. 2023; Köpf et al. 2023) predominantly consist of short samples.

In this work, we found that a simple and cheap approach which leverages a pre-built large and diverse short-prompt dataset works surprisingly well on long-context benchmarks. Specifically, we take the RLHF dataset used in Llama 2 Chat and augment it with synthetic self-instruct (Wang et al. 2022) long data generated by Llama 2 Chat itself, in the hope that the model can learn a diverse set of skills through the large amount of RLHF data and transfer that knowledge to long-context scenarios via self-instruct data. The data generation process focuses on QA-format tasks: starting from a long document in our pretraining corpus, we select a random chunk and prompt Llama 2 Chat to write question-answer pairs based on information in the text chunk. We collect both long and short form answers with different prompts. After that, we also adopt a self-critique step where we prompt Llama 2 Chat to verify the model-generated answers. Given a generated QA pair, we use the original long document (truncated to fit the model’s maximum context length) as the context to construct a training instance.

For short instruction data, we concatenate them as 16,384-token sequences. For long instruction data, we add padding tokens on the right so that models can process each long instance individually without truncation. While standard instruction tuning only calculates loss on the output tokens, we find it particularly beneficial to also calculate the language modeling loss on the long input prompts, which gives consistent improvements on downstream tasks (Section 4.3).

Main Results

To make long-context LLMs universally useful, an important desiderata is to ensure robust performance on standard short-context tasks. We verify our models’ performance on a series of common benchmarks following the previous work (Touvron et al. 2023). The aggregated results are shown in Table 1. Overall, we observe on-par and, in most cases, stronger results than Llama 2. Notably, we observe significantly improved results on coding, math, and knowledge intensive tasks such as MMLU. As shown in Table 2, our model outperforms GPT-3.5 on MMLU and GSM8k. This is in contrast to a previous work (Chen et al. 2023) which observes degradation on short tasks. We attribute the improvements to additional computation FLOPs and the knowledge learned from newly introduced long data.

Long Tasks

Different from previous works (Chen et al. 2023; Mohtashami and Jaggi 2023) that mostly rely on perplexity and synthetic tasks to gauge long-context performance, we perform long-context evaluation using real-world language tasks. We evaluate 0-shot performance on NarrativeQA (Kočiský et al. 2018), 2-shot on QuALITY (Pang et al. 2022) and Qasper (Dasigi et al. 2021), and 1-shot on QMSum (Zhong et al. 2021). The number of shots are decided based on the average sample length of each dataset (i.e., samples in Qasper and QuALITY are often much shorter than those of NarrativeQA). We focus these QA-style tasks because of the ease of prompt engineering We use simple prompt “{context} Q: {question}, A:” to evaluate all pretrained models. and less biased automatic evaluations. The input prompts are truncated from the left side if the prompts exceed the maximum input length of the model or 16,384 tokens. We compare with open-source long-context models available in Huggingface Transformers, namely Focused Transformer (Tworkowski et al. 2023a), YaRN (Peng et al. 2023), Xgen (Nijkamp et al. 2023), MPT (MosaicML 2023b; MosaicML 2023a) and Together’s Llama 2 fork (Together 2023). As shown in Table 3, our models achieve superior performance compared to these models. At the 7B scale, only “Together-7B-32k” can match our model’s performance. Note that this model is not a purely self-supervised model and has been finetuned using a large supervised dataset to improve its few-shot results. As the 7/13B variants of our models have been trained with 32k-token sequences, we also perform comparisons using 32,768 maximum prompts lengths and the results are consistent, as shown in Table 13.

Effective Context Utilization

To validate that our models can effectively use increased context window, we first show in Figure 2 that the results on each long task improve monotonically as we increase the context lengths. Inspired by (Kaplan et al. 2020; Hoffmann et al. 2022), we also found that the language modeling loss of our model follows a power-law plus constant scaling relationship with the context length (Figure 1), suggesting:

Our model continues to show gain in performance (on the language modeling loss) up to 32,768 tokens of text, despite having diminishing returns. Taking our 70B model for example, if we double the context length, we can expect the loss to be reduced by a factor of 2−β≈0.72^{-\beta}\approx 0.7 plus a model specific constant (1−2−β)⋅γ(1-2^{-\beta})\cdot\gamma.

Larger models can leverage the contexts more effectively, indicated by the larger β\beta value of the curves.

2 Instruction Tuning Results

We test our instruction tuned model on ZeroSCROLLS (Shaham et al. 2023) which bundles 10 long-context datasets spanning from summarization, question answering, to multi-document aggregation tasks. For a fair comparison, we use the same configuration (prompts, truncation strategy, and maximum generation lengths, etc) as specified by the benchmark. As shown in Table 4, without using any human annotated long context data, our 70B chat model is able to outperform gpt-3.5-turbo-16k on 7 out of the 10 tasks. In addition, we run evaluations on six new long tasks introduced in L-Eval (An et al. 2023) and again observe strong results, as shown in Table 17 in the Appendix. We see that the finetuned model is particularly good at QA tasks which is the main theme of the self-instruct data. We expect the performance to be further improved if more diverse data are used for finetuning.

It is worth mentioning that evaluating long-context LLMs is a nontrivial task. The automatic metrics used in these benchmarks are limited in many ways. For instance, the summarization tasks only come with a single ground-truth summary and the nn-gram matching metrics do not necessarily align with human preference. For QA and aggregation tasks, where the metric is less of a concern, truncating the input context might also remove the information necessary to answer the question. Another important caveat is that most proprietary models do not share their training data details, which makes it hard to take into consideration the potential leakage during public benchmark evaluation.

3 Human Evaluation

Complementary to the automatic evaluation benchmark results, we conduct human evaluations by asking annotators whether they prefer the generation from our instruction finetuned model or from proprietary models like MPT-30B-chat, GPT-4, GPT-3.5-turbo-16k, and Claude-2 in terms of helpfulness, honesty, and harmlessness. Unlike automatic metrics, humans are better at evaluating the quality of model responses for long context models because of the large space of acceptable answers. We focus on two major application scenarios with a total of 2,352 examples. For multi-turn conversation data, each prompt is a chat history based on which the model needs to generate a coherent response. For the multi-document search query answering application, the model is provided with a few most relevant documents retrieved from a search session and the corresponding search query. We then evaluate how well these models can leverage the information (retrieved documents) to answer the given query. Each comparison example was evaluated by 3 different human annotators. The standard win rate of our our model over each model is calculated by averaging the result of each comparison example and the final score along with the 95% confidence interval is shown in Figure 3. With very little instruction data, our model can achieve competitive performance against MPT-30B-chat, GPT-3.5-turbo-16k, and Claude-2. It is worth noting that human evaluation on longer context tasks is challenging and generally requires well trained and skilled annotators. We hope this study can not only give a sense of the potential of our instruction finetuned model on some long context downstream applications but also motivate future efforts in developing more robust long context automatic evaluations.

Analysis

In this section. We perform ablation experiments to justify our design choices (i.e. architecture modification, data mixes, and training curriculum) and quantify their contributions to the final performance.

Another interesting observation from the visualization is that RoPE introduces large “oscillation” in the long-range regions, which could be undesirable for language modeling (Sun et al. 2022). To investigate whether this effect hurts performance, we also explored another recently proposed variant of rotary encoding, xPos (Sun et al. 2022), which smooths the high-frequency component. Note that xPos with the default parameters suffers from the same decaying issue as RoPE and therefore, we also applied a similar decay fix to xPos.

Specifically, we empirically compare the following methods: the RoPE baseline, PI, our proposed RoPE with adjusted base frequency (denoted as RoPE ABF), and xPos ABF (visual comparisons in Figure 4). We report results on 1) long-sequence validation perplexity in Table 5 and Figure 5(a), 2) the first-sentence-retrieval context probing task We also test on the PassKey task as used in (Mohtashami and Jaggi 2023). All the model variants except RoPE can achieve perfect accuracy. We believe this task is overly simple for context probing. in Figure 5(b), and 3) some representative regular context tasks in Table 6 (to validate that long models do not degenerate on short-context tasks). All model variants are continually pretrained from the 7B Llama 2 checkpoint with additional 80B tokens organized as 32,768-token long sequences.

Overall, results on these evaluations suggest that RoPE ABF performs the best among all explored variants. In particular, we see that RoPE ABF is the only variant that can maintain its performance up to the full 32,768-token context window on the first-sentence-retrieval task. We also found that xPos ABF with less oscillation does not lead to substantial gains, suggesting that these artifacts are not detrimental to language modeling. While xPos is claimed to possess better extrapolation property (Sun et al. 2022), we found that, with the base frequency modification, xPos does not extrapolate better than RoPE (see Appendix C). In addition to empirical results, we provide a theoretical analysis of RoPE ABF and its difference to PI in Appendix B. We argue that RoPE ABF distributes the embedded vectors with an increased granularity when compared to RoPE PI, making it a easier for the model to distinguish between positions. It is worth noting that the relative distance between the embedded vectors has a linear dependence on the key parameter of RoPE PI and a logarithmic dependence on the key parameter of RoPE ABF, which coincides with our empirical observation that the base-frequency is not very sensitive and can be easily adjusted based on the max sequence length.

2 Pretraining Data Mix

The data used to continually pretrain our model combines existing datasets used by Llama 2 and new long text data. We also adjusted the data source mix ratio to up-weight long data samples. Our early experiments with 7B models confirms the significant improvements using this data mix for long-context tasks, as shown in the first two rows of Table 7. In this section, we aim to rigorously investigate the source of improvements. In particular, we are interested in differentiating the effects of the data length distribution and the quality of the corpus itself.

We perform two additional ablations using Llama 2’s pretrain datasets: 1) we remove the long text data from the Llama 2 dataset and continually pretrain our model with mostly short documents; 2) we increase the sample weights of existing long text data to be similar to the long text ratio used by proposed new model. Interestingly, even with most of the long texts removed, the model can still obtain most of the performance gain over Llama 2. We also find that there is no clear and consistent advantage as we greatly increase the long data ratio (the third row v.s. the fourth row in Table 7 and Table 8). We observe similar results on the first-sentence-retrieval task as shown by Figure 7 in the Appendix.

Based on the above ablations, we can see that adjusting the length distribution of the pretrain data does not provide major benefits. However, as we evaluate these model variants’ performance on standard short-context tasks, we find that new data mix also leads to large improvements in many cases, especially knowledge-intensive tasks like MMLU, as shown in Table 8. These results suggest that long-context LLMs can be effectively trained even with very limited long data and the improvements of our pretrain data over the one used by Llama 2 mostly come from the quality of the data itself, instead of the length distribution difference.

3 Instruction Tuning

We explored various strategies to instruction-finetune the pre-trained long context model which do not require any supervised long data. We start with only finetuning the models with short instruction data from Llama 2 Chat (referred as "RLHF V5" in (Touvron et al. 2023)) and then blend in with some pretrain data to avoid forgetting of previous long context continual pretraining. As demonstrated in Table 9, using only short instruction data can already produce a decent long model that significantly outperforms Llama 2 on various long-context tasks. On top of this dataset that only includes short prompts, we see that adding pretrain data (calculating language modeling loss on the whole sequence) can further boost the performance on most datasets. Inspired by this, we add the LM loss over the long context inputs when we finetune with self-instruct data. This simple trick makes learning more stable when we have unbalanced input and output lengths In our cases, the output lengths of most samples are a lot shorter than the those of the long-context inputs., which gives significant improvements on most of the tested tasks (the last two rows of Table 9).

4 Training Curriculum

Continual pretraining has demonstrated its efficacy in our experiments, but an open question still remains: does pretraining from scratch with long sequences yield better performance than continual pretraining? In this section, we study different training curricula and try to investigate if continual pretraining can offer competitive performance with less computation budget. We start off by pretraining a 7B model with 32,768 sequence length from start to the end. Then we explored various two-stage training curricula where we begin with 4096 sequence length and switch to 32,768 when the model completes 20%, 40%, 80% of whole training process. For all cases, we keep the same number of total training tokens and make sure the number of tokens per each gradient update remains constant (4 million tokens) by adjusting the batch size and sequence length accordingly.

We evaluate our models on the long-text QA tasks used in Section 4.2 and report the final models’ perplexity on different validation corpora. As shown in Table 10 and Table 11, continual pretraining from short context models can easily save around 40% FLOPs while imposing almost no loss on performance. These results also align with the training loss curves we observed from each run in Figure 6 – the models can quickly adapt to the increased sequence length and get to similar loss scale.

AI Safety

Despite showing excellent performance on various of downstream tasks, large language models are prone to generating harmful, misinformative, and biased contents (Lin et al. 2021; Hartvigsen et al. 2022; Dhamala et al. 2021; Ji et al. 2023). Long-context language models can process extended inputs in their context window, but at the same time, they also face a higher risk of jailbreak, especially through means such as prompt injection (Greshake et al. 2023). In this section, we evaluate the safety capability of instruction fine-tuned model using three standard academic benchmarks including TruthfulQA (Lin et al. 2021), ToxiGen (Hartvigsen et al. 2022), and BOLD (Dhamala et al. 2021), similar to (Touvron et al. 2023). We focus on the largest instruction fine-tuned model variant (i.e., 70B) and compare its results with both open sourced LLMs (Falcon-instruct Almazrouei et al. 2023, MPT-instruct MosaicML 2023a) and propriety LLMS (GPT-3.5, GPT-4 (OpenAI 2023), Claude-2 (Anthropic 2023)) in Table 12.

We observe that in general instruction fine-tuned model maintains similar safety performance compared to Llama 2 Chat and is safer and less biased compared to other open-source LLMs such as Falcon-instruct and MPT-instruct. AI safety is a complex domain and it can be extremely difficult to comprehensively evaluate all safety aspects of instruction fine-tuned model with three benchmarks. However, we hope our analysis can serve as a pilot study and provide directional signals on long-context large language models’ safety performance, which are not discussed in other works on the same topic (Tworkowski et al. 2023b; Ding et al. 2023; Chen et al. 2023). Currently the community also lacks dedicated safety benchmarks for long-context large language model evaluation and we plan to invest in this direction in our future work.

We evaluate instruction fine-tuned model on TruthfulQA (Lin et al. 2021) to benchmark its factuality. The benchmark consists of 817 questions covering 38 categories including health, law, finance, and politics (Lin et al. 2021). Similar to (Touvron et al. 2023), we use few-shot prompts with 6 random QA pairs for generation and then leverage two fine-tuned GPT-3 models to classify whether the generation is truthful and informative. We report the percentage of generations that are both truthful and informative as the final metric in Table 12.

ToxiGen

We measure the toxicity of instruction fine-tuned model using ToxiGen (Hartvigsen et al. 2022) where we check the percentage of toxic and hateful generations against 13 minority groups. Following (Touvron et al. 2023), we filtered out prompts where annotators disagree with each other on the target demographic group. We use the default ToxiGen classifier fine-tuned based on RoBERTa (Liu et al. 2019) to evaluate the level of toxicity of the model’s outputs. We report the percentage of toxic generations across all groups in Table 12.

BOLD

Bias in Open-Ended Language Dataset (BOLD) Dhamala et al. 2021 is used in this work to quantify how biased the models are against people from different demographic groups. This dataset consists of 23,679 prompts extracted from English Wikipedia covering five domains including race, gender, religion, political ideology and profession with 43 subgroups in total. Following Touvron et al. 2023, we exclude prompts belonging to Hinduism and Atheism religious subgroups as they only feature 12 and 29 prompts, respectively. After generations are inferred from each model, we leverage the Valence Aware Dictionary and Sentiment Reasoner (VADER) Hutto and Gilbert 2014 to perform sentiment analysis with a score ranging between -1 and 1. A positive score corresponds to a positive sentiment towards the subgroup mentioned in the prompt and vice versa. A sentiment score close to 0 indicates neutral sentiment which is desired. We report the average sentiment score across 43 demographic subgroups as the final metric for BOLD in Table 12.

2 Red Teaming Exercises

Currently there is no open-sourced safety benchmark designed for long-context understanding. To ensure that the models are safe in long context use scenarios, we performed internal red teaming to better understand the vulnerability of our chat model. We attack the model by feeding long contexts (e.g., long conversations) to it, followed by adversarial prompts covering risky areas including illicit and criminal conducts (e.g., terrorism, theft, and human trafficking), hateful and harmful behaviors (e.g., defamation, self-harm, eating disorders, and discrimination), and unqualified advice Touvron et al. 2023. Through manual inspection, we did not observe significant risks compared to Llama 2 Chat Touvron et al. 2023. We plan to invest more in new attack vectors against long context large models in future work.

Limitations

The our model proposed in this paper has not yet been finetuned for a wide range of long-context applications, such as creative writing that require long-form outputs. Applying existing alignment recipes, e.g., RLHF, for various scenarios is expensive and nontrivial. Even skilled annotators may struggle to the intricate details in dense texts. In this regard, we consider developing efficient alignment methods for long LLMs to be a very valuable direction for future research.

Tokenizer Efficiency.

While the proposed our model series can consume contexts up to 32,768 tokens, the actually number of words our model can take is largely affected by the tokenizer behaviour. The tokenizer used by the Llama series has a relatively small vocabulary (32k symbols) and often produces longer sequences compare to the sequences given by GPT-3.5’s tokenizer – we observe our tokenizer often produce 10%10\% more tokens on average. Additionally, the tokenizer we use also cannot efficiently handle whitespace, making it inefficient to process long code data.

Hallucination.

Like other LLMs, we have observed hallucination issue when testing the proposed our model. While this issue is common for short-context models, tackling with this problem for long-context models can be more pronounced because of the dense information they consume and the insufficient alignment process.

Conclusion

We present a series of long-context LLMs that leverage a simple yet necessary position encoding refinement and continual pretraining to achieve strong long-context performance. Our long context scaling is performed by continually pretraining from Llama 2 with additional 400B tokens and outperform Llama 2 on both short and long-context tasks. Our models also demonstrate superior performance compared to existing open-source long-context models and compare favorably against gpt-3.5-turbo-16k on a suite of long-context tasks after a simple instruction finetuning procedure without human supervision. We complement our results with a comprehensive analysis, providing insights on the influences of various factors including the nuances of position encodings, the data mix, and the pretraining curriculum on the final performance. We hope our study could make long-context LLMs more accessible and facilitate further advancements in this field.

Acknowledgement

We would like to thank Nikolay Bashlykov, Matt Wilde, Wenyin Fu, Jianyu Huang, Jenya Lee, Mathew Oldham, and Shawn Xu for their invaluable support on the data, infrastructure, and various other aspects of this project.

References

Appendix A More Results

Appendix B Theoretical Analysis of Positional Encodings

The purpose of this mapping is to help the attention module to separate the vectors corresponding to two instances of the same token that are situated at different positions in the input sequence.

Aiming at extending the sequence length of a transformer pretrained with a particular positional embedding ff from LL to L^,\hat{L}, we would like to come up with a positional embedding f^\hat{f} that minimizes the distance between the old and the new images of the embedded vectors:

With this in mind, we consider two different methods to extend the sequence length of a trained transformer: Position Interpolation (PI) parameterized with α\alpha, and Adjusted Base Frequency (ABF) parameterized with β.\beta. These two methods correspond to the following embedding curves:

This leaves us with a multi-objective decision selecting the positional embedding for a model with extended context: on one hand, f^\hat{f} should be chosen so that it minimizes d(f,f^),d(f,\hat{f}), while on the other hand its value of q(f^)q(\hat{f}) should be big enough.

The red dots on the line correspond to 1111 integer values of t.t.

Figure 8(b) aims to illustrate the impact of Position Interpolation on the relative position of the mapped vectors. The distance between the consecutive points got reduced considerably compered to Figure 8(a). The impact of Adjusted Base Frequency is illustrated on Figure 8(c). The distance between the consecutive points remained almost the same as on Figure 8(a), although the minimal distance between points got considerably reduced due to the increased frequency of the helix. This effect of increased frequency of the helix would be reduced in the high dimension setting. The value of the coefficient aa for the helix depicted on Figure 8(a) is two times larger than the value of the coefficient aa for the helix depicted on Figure 8(c). If the dimension of the input of the attention mechanism is d=128,d=128, then the difference between θ1=b−2jd\theta_{1}=b^{-\frac{2j}{d}} at b=10,000b=10,000 and θ1=b−2jd\theta_{1}=b^{-\frac{2j}{d}} at b=500,000b=500,000 is only 6%.6\%. Thus, we further focus specifically on the distance between the consecutive images of the embeddings.

We make a formal comparison between Positional Interpolation and Adjusted Base Frequency by analytically comparing the pairwise distances between the images given by fRoPE+PIf^{RoPE+PI} and fRoPE+ABFf^{RoPE+ABF} for consecutive integer values of tt. This corresponds to the evaluation of q(f^)q(\hat{f}) discussed earlier. We will measure the distance between embedding images in terms of the Euclidean sine similarity metric since all versions of RoPE are norm-preserving.

The following result states that in a high-dimensional space, the sine similarity sin⁡∠(fRoPE+ABF(x,n+1),fRoPE+ABF(x,n))\sin\angle(f^{RoPE+ABF}(x,n+1),f^{RoPE+ABF}(x,n)) between two consecutive embedding images of a vector xx can be bounded with a value proportional to (log⁡b+log⁡β)−1.(\log b+\log\beta)^{-1}. Moreover, the similarity sin⁡∠(fRoPE+PI(x,n+1),fRoPE+PI(x,n))\sin\angle(f^{RoPE+PI}(x,n+1),f^{RoPE+PI}(x,n)) can be bounded using α(log⁡b)−1.\alpha(\log b)^{-1}.

min⁡kxk2∥x∥2Cd≤sin⁡∠(f(x,n+1),f(x,n))≤max⁡kxk2∥x∥2Cd\frac{\min_{k}x_{k}^{2}}{\|x\|^{2}}C_{d}\leq\sin\angle(f(x,n+1),f(x,n))\leq\frac{\max_{k}x_{k}^{2}}{\|x\|^{2}}C_{d}

where lim⁡d→∞Cd≈{(log⁡b+log⁡β)−1 if f=fRoPE+ABFα(log⁡b)−1 if f=fRoPE+PI\lim_{d\to\infty}C_{d}\approx\begin{cases}(\log b+\log\beta)^{-1}\text{ if }f=f^{RoPE+ABF}\\ \alpha(\log b)^{-1}\text{ if }f=f^{RoPE+PI}\end{cases} under the assumptions of α≪1\alpha\ll 1 and b≫1.b\gg 1.

Let us begin the proof by writing down the expressions for the inner product between two images of RoPE variants.

⟨fRoPE+PI(x,m),fRoPE+PI(x,n)⟩=∑j=0d2−1(x2j2+x2j+12)eib−2jdα(m−n)\langle f^{RoPE+PI}(x,m),f^{RoPE+PI}(x,n)\rangle=\sum_{j=0}^{\frac{d}{2}-1}\left(x_{2j}^{2}+x_{2j+1}^{2}\right)e^{ib^{-\frac{2j}{d}}\alpha(m-n)}

⟨fRoPE+ABF(x,m),fRoPE+ABF(x,n)⟩=∑j=0d2−1(x2j2+x2j+12)eib−2jdβ−2jd(m−n)\langle f^{RoPE+ABF}(x,m),f^{RoPE+ABF}(x,n)\rangle=\sum_{j=0}^{\frac{d}{2}-1}\left(x_{2j}^{2}+x_{2j+1}^{2}\right)e^{ib^{-\frac{2j}{d}}\beta^{-\frac{2j}{d}}(m-n)}

From them, we can derive the expressions for the Euclidean sine similarity between the images of the positional embeddings:

sin⁡∠(fRoPE+PI(x,m),fRoPE+PI(x,n))=∑j=0d2−1(x2j2+x2j+12)sin⁡(b−2jdα(m−n))∑j=0d−1xj2\sin\angle(f^{RoPE+PI}(x,m),f^{RoPE+PI}(x,n))=\frac{\sum_{j=0}^{\frac{d}{2}-1}\left(x_{2j}^{2}+x_{2j+1}^{2}\right)\sin(b^{-\frac{2j}{d}}\alpha(m-n))}{\sum_{j=0}^{d-1}x_{j}^{2}}

sin⁡∠(fRoPE+ABF(x,m),fRoPE+ABF(x,n))=∑j=0d2−1(x2j2+x2j+12)sin⁡(b−2jdβ−2jd(m−n))∑j=0d−1xj2\sin\angle(f^{RoPE+ABF}(x,m),f^{RoPE+ABF}(x,n))=\frac{\sum_{j=0}^{\frac{d}{2}-1}\left(x_{2j}^{2}+x_{2j+1}^{2}\right)\sin(b^{-\frac{2j}{d}}\beta^{-\frac{2j}{d}}(m-n))}{\sum_{j=0}^{d-1}x_{j}^{2}}

Let’s put m=n+1m=n+1 to compare the distance between the two consecutive positional embedding images of the same vector x.x.

∥x∥2sin⁡∠(fRoPE+PI(x,n+1),fRoPE+PI(x,n))=∑j=0d2−1(x2j2+x2j+12)sin⁡(b−2jdα)\|x\|^{2}\sin\angle(f^{RoPE+PI}(x,n+1),f^{RoPE+PI}(x,n))=\sum_{j=0}^{\frac{d}{2}-1}\left(x_{2j}^{2}+x_{2j+1}^{2}\right)\sin(b^{-\frac{2j}{d}}\alpha)

∥x∥2sin⁡∠(fRoPE+ABF(x,n+1),fRoPE+ABF(x,n))=∑j=0d2−1(x2j2+x2j+12)sin⁡(b−2jdβ−2jd)\|x\|^{2}\sin\angle(f^{RoPE+ABF}(x,n+1),f^{RoPE+ABF}(x,n))=\sum_{j=0}^{\frac{d}{2}-1}\left(x_{2j}^{2}+x_{2j+1}^{2}\right)\sin(b^{-\frac{2j}{d}}\beta^{-\frac{2j}{d}}) Due to the range of b,αb,\alpha and β\beta that is typically considered, we can bound the arguments of the sine functions as 0<αb−2jd≤10<\alpha b^{-\frac{2j}{d}}\leq 1 as well as 0<(βb)−2jd≤1.0<(\beta b)^{-\frac{2j}{d}}\leq 1. Using that, we derive that sin⁡(b−2jdβ−2jd)\sin(b^{-\frac{2j}{d}}\beta^{-\frac{2j}{d}}) and sin⁡(b−2jdα)\sin(b^{-\frac{2j}{d}}\alpha) are non-negative as well as xj2x_{j}^{2} for any j∈{1,…d}.j\in\{1,\ldots d\}. Thus, the following inequalities hold:

Carrying min⁡kxk2\min_{k}x_{k}^{2} and max⁡kxk2\max_{k}x_{k}^{2} out of the summation signs, we obtain

Introducing CdABF=∑j=0d2−1sin⁡(b−2jdβ−2jd)C^{ABF}_{d}=\sum_{j=0}^{\frac{d}{2}-1}\sin(b^{-\frac{2j}{d}}\beta^{-\frac{2j}{d}}) and CdPI=∑j=0d2−1sin⁡(b−2jdα)C^{PI}_{d}=\sum_{j=0}^{\frac{d}{2}-1}\sin(b^{-\frac{2j}{d}}\alpha) proves the first part of the Theorem:

Now, considering the limit of Cd,C_{d}, we notice that due to the inequalities on the arguments of the sines, the following bounds hold:

Using the formula of geometric sums and a corollary of the exponential (second) foundational limit, we establish the limits of the sums of these bounds as d→∞d\to\infty:

Substituting these into the bounds on lim⁡d→∞Cd,\lim_{d\to\infty}C_{d}, one achieves:

From these bounds, one can see that in the setting considered within this paper, where b=10000b=10000 and α<1/4,\alpha<1/4, the approximation of lim⁡d→∞Cd\lim_{d\to\infty}C_{d} used in the statement of the Theorem is of a high quality.

Based on this theoretical derivation, we return to the interpretation of our experimental resuts. On one hand, the experiments have shown that the model can adapt to the new sequence length with both RoPE PI (α=1/4\alpha=1/4 or α=1/8\alpha=1/8) and RoPE ABF (β=50\beta=50). Thus, we can conclude that the chosen hyperparameters provide a sufficient degree of approximation of RoPE images under b=10000.b=10000. In other words, both d(f,fRoPE+ABF)d(f,f^{RoPE+ABF}) and d(f,fRoPE+PI)d(f,f^{RoPE+PI}) are small enough to allow rapid adaptation. On the other hand, comparing the expressions of CdC_{d} for RoPE ABF and RoPE PI, we can observe that for the values of α=14\alpha=\frac{1}{4} or α=18\alpha=\frac{1}{8} and b=10000b=10000 that were used in our experiments, the granularity (the distance between two consecutive images of RoPE) is much lower for the RoPE PI (α(log⁡b)−1≈0.027\alpha(\log b)^{-1}\approx 0.027) than for RoPE ABF ((log⁡b+log⁡β)−1≈0.076(\log b+\log\beta)^{-1}\approx 0.076) with β=50.\beta=50. We further hypothesise that the higher degree of granularity is related to the higher evaluation on the downstream tasks of the RoPE ABF variant compared to RoPE PI because it makes the task of distinguishing between the positional embedding images simpler for the model. In other words, this corresponds to the case of q(fRoPE+ABF)>q(fRoPE+PI).q(f^{RoPE+ABF})>q(f^{RoPE+PI}).

Throughout this consideration we implicitly assumed that the distance between the consecutive images of an embedding is smaller than the distance between any other pair of the images. While this assumption is likely to hold true in a high-dimensional space, significantly increasing the parameter of β\beta in RoPE ABF may violate this assumption due to the changed geometry of the embedding curve.

Appendix C Length Extrapolation Results

Despite not the focus of this work, extrapolation is an important property for long context models. Extrapolation refers to a model’s ability to conduct inference on input sequences that are longer than its training sequences. We evaluate how our 70B model extrapolates with two tasks:

Validation loss at each position: In Figure 9(a), we visualize the average loss at each position of the 32,768 sequence length where the first 16,384 is the interpolation area (within training sequence length) and the second half is extrapolation. We use 50 batches of samples and average across them. To make plots smoother, we also take the mean of losses every 500 positions. As we can see, our 70B model with either RoPE ABF or xPos ABF maintain the loss in the extrapolation area. To contrast this, we also plot the result for Llama 2 with 4,096 context window: the loss explodes after the position goes beyond training sequence length, which suggests that Llama 2 does not extrapolate effectively.

Synthetic first-sentence-retrieval task: To complement validation loss evaluation, we also test our 70B model with two different PEs on the context probing task. Unlike validation loss task where it is hard to find data samples that require very long range dependencies consistently, first-sentence-retrieval imposes a very strict requirement for models to attend with a specific length. In Figure 9(b), we visualize the results up to 32,768 where we do see some performance degradation when the model needs to extrapolate. In addition, we observe that, despite often considered as having better extrapolation properties, xPos ABF does not outperform RoPE ABF in our setting.

Appendix D Self-Instruct Data

As described in Section 4.3, we use Llama 2 Chat to bootstrap self-instruct data for instruct finetuning. In this section we describe the detailed procedure as well as providing the necessary prompts used for generating this dataset. The main challenge is that we need an automated process of generating long context instruct data with only short context models at hand. The core idea behind this is to split the long documents into chunks of texts that can fit into short model’s context and apply self-instruct. We focus primarily on question answering dataset. We first split the long document into smaller chunks, and for each chunk we construct a prompt as in Figure 10 which gets fed into Llama 2 Chat to get a question-answer pair. To diversify the question types, we randomly choose between the two prompts that ask for either normal or short answers. Once we extract the question and answer from the response (using tags as required by the prompt), we can construct long question answering instruct data together with the original long document, using the templates in Figure 11 of the corresponding answer type.