Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
Yunpeng Huang, Jingwei Xu, Junyu Lai, Zixu Jiang, Taolue Chen, Zenan Li, Yuan Yao, Xiaoxing Ma, Lijuan Yang, Hao Chen, Shupeng Li, Penghao Zhao
Introduction
In recent years, leveraging techniques from deep learning , especially the surge of Transformer-based models like BERT , GPT and their variants , Natural Language Processing (NLP) has significantly advanced, empowering machines to understand and generate human language , thus revolutionizing numerous tasks in Natural Language Understanding (NLU) like sentiment analysis , Natural Language Generation (NLG) like document summarization , as well as other domains like computer vision and autonomous driving . Furthermore, in the wake of ChatGPT , PaLM , GPT4 , etc, the Transformer-based Large Language Models (LLMs) which scale up to 1B100B parameters to empower emergence abilities , have shown a new exhilarating path towards Artificial General Intelligence (AGI) , and been rapidly adopted in a myriad of human-interactive applications, like chatbots , programming assistants and educational tutors .
Transformer is an elaborate deep neural network model, which incorporates many great preceding designs and comprises diverse novel components to solve the sequence-to-sequence language modeling problem in machine translation at the very beginning . The contemporary LLMs mostly derive their foundation from the Transformer architecture by employing its full or part modules . Among those components, Transformer-based LLMs own the success mainly due to the core well-designed attention mechanism that captures global dependencies of each pair of tokens across the whole input, enabling the model to handle sequences with intricate relations. While the attention mechanism offers remarkable performance, its quadratic time and space complexities with respect to input sequence length lead to a significant computational resource bottleneck, which imposes limitations on not only the permissible input text length during training, but also the effective context window of prompts due to unsatisfactory efficiency and expensive caching memory consumption as the generated tokens increase during inference. To be worse for inference, the LLMs also suffer from performance degeneration when facing sequences longer than the ones in training, owning to poor generalizable mechanism design for input length.
However, with LLMs deeply ingrained in various applications that require long-context comprehension and generation , the demand for long-context LLMs capable of comprehending and generating extremely long sequences effectively and efficiently becomes increasingly indispensable and urgent. Consequently, researchers have devoted significant efforts to enhancing the Transformer architecture to address the long-context problem in LLMs, including optimization on the efficiency of attention (Sec.3), context window extension with extra memory mechanisms (Sec.4), effective length generalization with extrapolative positional embeddings (Sec.5), context pre/post-processing (Sec.6), and other miscellaneous methods (Sec.7) such as specific pre-training objectives, mixture of experts, quantization, parallelism, etc.
Existing surveys. The field of long-context LLMs has become one of the hottest and most rapidly developing research areas on LLMs recently, with some existing surveys summarizing the related work of literature. Among those, offers a cursory overview of long document summarization yet refrains from delving deeply into the intrinsic techniques of long text modeling. In and , they both primarily concentrate only on augmenting the computational efficiency of Transformers in long-text scenarios. Although underscores the challenges LLMs face when engaging with extensive sequences, its discussed methods predominantly align with efficient Transformers, akin to and . A more recent contribution bears the closest resemblance to our study, introducing approaches in long-text modeling and applications with Transformers, covering pre-processing techniques, part of efficient Transformers, and special characteristics of lengthy documents. However, there is still a lack of comprehensive studies to review the literature on the advancement in breaking the barriers of context length across all stages for more intricate and scalable Transformer-based LLMs by exploring the Transformer’s architecture from an operational perspective.
The objective of this survey is to thoroughly review the panorama of literature on architecture evolution for scaling the effective context window length of present Transformer-based LLMs from a methodological perspective. The key contributions are as follows:
We build a holistic taxonomy to categorize five pieces by breaking down the Transformer architecture and then delving into the existing methodologies in enhancing long-context LLMs during each stage, including pre-training, fine-tuning, inference, and pre/post-processing.
We explore the widely-used evaluation necessities, comprising datasets, metrics, and baseline specifically assessing the long-context capabilities of LLMs, followed by some popular toolkits to optimize LLMs’ efficiency and effectiveness for both training and inference, such as libraries, systems, and compilers.
We identify key challenges to revamping the Transformer structure for handling extensive contexts, with corresponding future directions to further push the frontier.
Considering the extremely rapid growth of this field and the survey might fall behind soon, we build a repository that gathers relevant literature within this specific domain. We shall keep it updated continuously to help readers keep pace with the newest advancements.
Organization of this survey. Sec.2 gives an overview of long-context LLMs, including the preliminaries about objectives and stages for language modeling and critical components of Transformer-based LLMs, the structure limitation analyses for LLMs to deal with lengthy contexts and the taxonomy of existing efforts on advancing Transformer architecture. Then, we mainly delve into the discussion of each part of methodologies from the taxonomy in next five sections 3, 4, 5, 6,7, corresponding to related modules in Transformer architecture. In Sec.8, we also summarize the necessities for long-context capabilities evaluation and collect some popular optimization toolkits to augment LLMs about effectiveness and efficiency during training and inference. Then in Sec.9, we explore the critical challenges and corresponding potential avenues lighted up by them, as well as draw insights from existing breakthroughs. Finally, Sec.10 closes this survey with overarching conclusions regarding the the panorama of this domain as well the motivation of this study.
Overview
In this section, we start with the preliminaries (Sec.2.1) from the fundamental language modeling objectives, typical modeling stages to critical architecture modules found in Transformer-based decoder-only LLMs, as depicted in Fig.1 (a). Subsequently, we offer brief analyses of the architecture limitations when LLMs encounter extensive context windows (Sec.2.2). Finally, we present a comprehensive methodological taxonomy (Sec.2.3) aimed at enhancing the long-context capabilities of LLMs through architectural innovations (see Fig.1 (b)). This taxonomy serves as a guideline for the next five sections 3, 4, 5, 6,7.
Language Modeling. The heart of LLMs lies the foundational task of language modeling, which aims to empower neural networks to comprehend and generate human language. From a mathematical perspective, the essence of neural language modeling is to approximate the precise log-probability of the occurrence of any given text, denoted as , where stands for the network parameters to learn, and comprises a sequence of elements representing natural language including words, punctuation, mathematical symbols and more. However, this task encounters a significant practical hurdle known as the curse of dimensionality, which stems from the exponential growth of possibilities as increases. To avoid the impracticality, LLMs employ variations mainly including masked language modeling (MLM) and causal language modeling (CLM). The former MLM is to predict masked tokens based on the bidirectional remaining unmasked tokens, whose objective can be written as Eq.1, that maximizes the conditional probability of the th token given all others, where denotes the index set of the masked tokens. In contrast, the objective of CLM is to predict the next token, i.e., maximize the conditional probability of each token, given the unidirectional preceding ones (see Eq.2). In this setup, casual LLMs can effectively leverage the temporal dependencies inherent in natural language sequences, enabling LLMs to generate coherent and contextually relevant text.
Modeling Stages. Nowadays, the operation of a LLM often undergoes a multi-stage modeling process. Initially, during the pre-processing stage, raw text data is segmented and tokenized into individual (sub)words called tokens predefined in a vocabulary, using algorithms like BPE . Then, in the pre-training stage, the model is trained on vast text corpora, with the object of either MLM or CLM, to capture semantic patterns and linguistic structures of natural language. Once pretrained, the model proceeds to the fine-tuning stage, where it is further trained with a few epochs on task-specific data with extra heads to learn sometimes. Finally, the fine-tuned model is deployed into downstream scenarios to predict expected answers in an inference mode. Particularly, the casual LLMs are pre-trained and fine-tuned with the same CLM objective but a different corpus. During the inference step, the model predicts from the probability distribution of the vocabulary by some decoding strategy such as greedy search, beam search, nucleus sampling , to generate contextually coherent responses to prompts in a token-by-token autoregressive paradigm.
Decoder Block. The vanilla Transformer architecture proposed in mainly comprises an Encoder and a Decoder, each stacked with multiple identical blocks. The skeleton of each block is mostly compatible with the one portrayed in Fig.1 (a). In general, the first block will take the tokenized sequence encoded by a word embedding layer, following a multi-head scaled-dot self-attention (MHA) layer with an attention mask corresponding to specific language modeling objectives and a feed-forward network (FFN) layer. Both the MHA and FFN layers are enriched with layer normalization and residual connections at every entrance/exit of the block. Then, each higher-level block takes the output hidden states from the previous block as input, represents them with its MHA and FFN layers, and feeds them to the next block. The final output hidden state from the last block is fed into a linear layer called language modeling head, and the output logits will be transformed into a probability distribution over the target vocabulary through softmax operation. Notably, the slight difference between the Encoder block and Decoder block in an Encoder-Decoder Transformer is that the latter additionally interfaces with the Encoder’s output via an cross-attention (CA) layer before feeding into the FFN layer.
However, such a binary structure was originally designed for sequence-to-sequence modeling in machine translation tasks. Subsequently, it has given rise to several variations aimed at more general language modeling objectives like MLM and CLM. The BERT series harnesses only the Encoder with MLM to enhance bidirectional information, serving as a discriminative model. Conversely, the GPT series utilizes only the Decoder with CLM, focusing on unidirectional generative models. T5 and BART variants, however, treat each NLP task as a text-to-text conversion, leveraging both Encoder and Decoder. The decoder-only generative model architecture has recently become the predominant choice for current LLMs. Notable examples include GPT4 , PaLM , LLaMA , and GLM , among others.
It is worth noting that SinPEs are initially applied on the word embeddings before entering the Encoder or Decoder blocks by addition. In contrast, as shown in Fig.1 (a), RoPEs are applied to in each attention layer before the kernel operations by equivalent element-wise vector multiplication to save registered buffer memory.
2 Limitation Analyses
Attention Complexity. In typical scenarios where , the computational complexity of MHA can be concisely summarized as follows: It involves time complexity, comprising for QKV projection, for the computation of , for the softmax operation to obtain , for the multiplication of and , and for the output projection of . And it incurs space complexity, involving for embeddings of , and additional buffers for storing weights and . Consequently, both temporal and spatial computational costs exhibit a quadratic increase with the expansion of the sequence length, which can be burdensome for both training and inference.
In-context Memory. LLMs lack an explicit memory mechanism, relying solely on the KV cache to store representations of all previous tokens in a list. This design implies that once querying is completed in one call, the Transformer does not retain or recall any previous states or sequences in subsequent calls unless the entire history is reloaded token by token into the KV cache. Consequently, the Transformer possesses only an in-context working memory during each call, as opposed to an inherent memory mechanism such as Long Short-Term Memory (LSTM) . This statelessness offers computational advantages in terms of parallelism but presents challenges in tasks like chatbot applications , where long-term memory retention is essential.
Max-Length Constraint. During the training phase, engineers typically need to determine a crucial hyperparameter max-length, denoted as throughout this paper. This hyperparameter represents the upper bound on sequence length for any training sample in a batch. It is commonly set to values such as 1k, 2k, or 4k based on the available computational resources to avoid Out-of-Memory (OOM) errors on GPUs. However, during inference, LLMs service providers must also either restrict the length of user prompts or automatically truncate them to align with the predefined , even though inference resources are typically more abundant than during training. Note that none of the Transformer modules inherently require such restrictions since all learned weights depend solely on dimension sizes. So, theoretically, Transformers can process sequences of any length as long as the resources are sufficient. Unfortunately, current Language Models have shown noticeable performance degradation when handling input sequences exceeding , often resulting in repetitive and implausible outputs.
3 Taxonomy
Building upon the foundational insights presented in Sec.2.1 and the limitations discussed in Sec.2.2, there are multiple avenues to explore for advancing the Transformer structure to endow LLMs with long-context capabilities, such as reducing attention complexity during training, designing efficient memory mechanisms, and enhancing the ability for length extrapolation, as outlined in , where the model is trained on short sequences but tested on longer ones during inference. Consequently, in this paper, we provide a comprehensive review of recent advancements in methodologies aimed at improving the long-context capabilities of LLMs throughout various stages, and we organize them into a unified taxonomy, as illustrated in Fig.1 (b). Specifically, these methods are categorized into five main clusters as follows:
Efficient Attention (Sec.3): these methods focus on implementing efficient attention mechanisms with reduced computational demands, even achieving linear complexity. By doing so, they enable the extension of the effective context length boundary of LLMs during inference by directly increasing in the pre-training stage.
Long-Term Memory (Sec.4): to address the limitations of the in-context working memory, some approaches aim to design explicit memory mechanisms that compensate for the lack of efficient and effective long-term memory in LLMs.
Extrapolative PEs (Sec.5): recent efforts have been made to enhance the length generalization capability of LLMs by improving the extrapolative properties of existing positional encoding schemes.
Context Processing (Sec.6): in addition to methods that enhance specific low-level Transformer modules, some approaches involve wrapping off-the-shelf LLMs with additional context pre/post-processing. These methods ensure that the input fed to LLMs in each call always meets the maximum length requirement and breaks the context window limit by introducing multiple calling overheads.
Miscellaneous (Sec.7): this section explores various general and valuable methods that do not neatly fit into the previous four categories, offering a broader perspective on advancing long-context capabilities in LLMs.
Efficient Attention
The first category of methods is dedicated to optimizing attention mechanisms, especially focusing on the kernel operations that make the module the computational bottleneck of the Transformer (see Eq.4). This approach enables the expansion of the effective context length boundary for LLMs during inference by directly increasing the hyperparameter in the pre-training stage. We further categorize these methods into five distinct strategies, each with a specific focus: Local Attention (Sec.3.1), Hierarchical Attention (Sec.3.2), Sparse Attention (Sec.3.3), Approximated Attention (Sec.3.4), and IO-Aware Attention (Sec.3.5).
The traditional attention mechanism is characterized by its global and full attention nature, wherein every token is expected to attend to every other token, resulting in quadratic time and space complexities. Considering the significance of local context in certain applications , various approaches have been introduced to implement local attention mechanisms in recent years. These mechanisms restrict each token’s attention to its neighboring tokens instead of all tokens, and the variations among these approaches arise from the heuristic criteria to determine who qualifies as a token’s neighbor, as depicted in Fig.2.
Block-wise Attention. One straightforward approach to implementing local attention involves segmenting the input sequence into non-overlapping blocks. As proposed in BlockBERT , within each block of fixed size , tokens are only allowed to attend to other tokens within the same block. This block-wise attention involves performing full attention calculations within each block for iterations, resulting in a time complexity of and a memory complexity of . However, this approach restricts the global receptive field and can limit the ability to model long-term dependencies. To address this limitation, Bi-BloSAN introduces an inter-block attention mechanism to capture long-range dependencies. Sinkhorn employs a differentiable ranking network to sort blocks, allowing each token to attend to tokens in the newly sorted block, thus enabling a quasi-global receptive field. SPADE augments state space models (SSMs) to address long-range dependency limitations. Additionally, Landmark Attention introduces a new token called the landmark token for each block, enabling block-wise representations by training the attention mechanism to use it to select relevant blocks. In the fine-tuning stage, LongLoRA introduces shift short attention (S2-Attn) on top of LoRA , shifting tokens by half the block size in half of the attention heads to ensure information flow between neighboring blocks.
Sliding Window Attention. Inspired by convolutional neural networks (CNNs) , another approach is to use sliding-window techniques, as demonstrated in Longformer . In this method, each token is assigned a consecutive fixed window and is allowed to attend only to the previous adjacent tokens as its neighbors. To extend the receptive field similar to dilated convolution , the window is dilated with gaps of size dilation , enabling each token to attend to tokens as far as away. To aggregate global information without additional computation, global attention is also applied to a few pre-selected positions where special tokens like [CLS] are located, decreasing computation complexity to .
Global-Local Hybrid Attention. A similar global-local attention mechanism has also been adopted in ETC and LongT5 , which explicitly or implicitly construct auxiliary global tokens to represent the segment information with global attention while only applying local attention to tokens in the source. We can also consider it as a hierarchical organization of attention receptive fields related to Sec.3.2. Another intriguing technique of the global token comes from the very recent streamLLM , where they observe an interesting phenomenon, namely attention sink, that not only keeping the KV of initial tokens during inference will largely recover the performance of sliding window attention, but adding a placeholder token during pre-training can further improve streaming deployment as well. They demonstrate that the emergence of attention sink is due to the strong attention scores towards initial tokens as a “sink” even if they are not semantically important. Similar strategies to keep attending to the starting tokens have also been proposed in the very recent Lm-infinite , where they propose a shaped mask and a positional distance constraint.
2 Hierarchical Attention
To further think of either the global token techniques , or the inter-block attention mentioned above, we can regard them as introducing some hierarchical features to self-attention to compensate with more global information from the higher-level attention while keeping the low computation cost from the low-level local attention at the same time. From this view, many works have explored various hierarchical mechanisms that introduce a structured hierarchy into self-attention, leveraging both higher-level global information and lower-level local attention for multi-scaled contextual receptive fields.
Two-Level Hierarchy. HAN pioneers the use of a two-level attention mechanism. It first applies self-attention to word features to obtain a sentence representation, and then employs self-attention on sentence-level features to generate document-level features. This hierarchical approach improves efficiency and performance in document classification tasks. Subsequently, similar hierarchical attention mechanisms have led to significant advancements in various document-level tasks, including machine translation and document summarization .
Multi-Level Hierarchy. In contrast to the typical binary level structure above, BPT introduces a more elaborated fine-to-coarse attention mechanism that operates on multi-scale spans through binary partitioning. Token nodes can attend to smaller-scale spans for closer context and larger-scale spans for more distant context. This approach formalizes the hierarchical structure as a graph neural network and updates it using graph self-attention . A simpler variation is seen in Adaptive Span Transformer , which employs a soft attention masking function to non-increasingly map relative distances to real values in the range $$. This function controls the span of attention for each head, allowing the model to attend to different context spans.
Building upon the hypothesis, the attention matrix in many NLP tasks holds a hierarchical low-rank structure, H-Transformer-1D introduces hierarchical attention that partitions the attention matrix into different blocks with varying low-rank ranges, enabling different levels of approximation. This approach reduces the overall runtime and memory cost complexity to a linear scale of , with the number of hierarchy levels denoted as . Viewing full-attention as a conditional expectation over embeddings at each location, Combiner approximates this conditional distribution with structured factorization on token regions. Tokens can then attend to others either directly or through indirect attention to abstractions, which are conditional expectations from corresponding factorized local regions. This approach also leverages sparse attention patterns, as discussed in the next Sec.3.3, to provide sub-quadratic low computation and memory complexity while maintaining full-attention expressiveness.
Generally speaking, hierarchical attention mechanisms derive from the same principles of contextual locality present in natural languages as local attention. However, they incorporate a more elaborated structure, often designed heuristically, to strike a balance between capturing long-range contextual dependencies and maintaining low-level computational efficiency.
3 Sparse Attention
While some approaches have introduced heuristics for achieving locality and hierarchical structure within self-attention, another direction explores the sparsity patterns inherent in full attention matrices . These methods aim to introduce a sparse attention mask, denoted as , where each row assigns a sparse set of indices that the -th token attends to. These sparsity-based attention mechanisms offer both computational efficiency and the ability to capture global context information. Figure 3 provides a visualization of these sparse attention mechanisms.
Adaptive Sparsity Patterns. Instead of fixed sparse indices set only dependent on locations, some approaches seek sparsity adaptively in a learnable manner, taking into account embedding values. Expire-Span introduces a learnable scalar in the range $Q,KO(L\sqrt{L}d)\rhoO((1-\rho^{2})L^{2}d)$.
Graph Sparsification. Furthermore, some other works treat full attention as a fully connected graph, with nodes representing embeddings of each token and edges denoting connections through attention. These approaches frame sparsity as a graph sparsification problem. For instance, Star-Transformer introduces a star-shaped topology, where each satellite node attends to local neighbors with a ring connection and a virtual relay node with the radial connection. In contrast, BigBird incorporates sparsity based on random graph theory, allowing each query to attend to a random number of keys with a fixed probability. It also absorbs sliding-window local attention with window size and global token techniques in its design. This approach reduces the quadratic dependency to linear complexity, specifically , stacked with three efficient attention mechanisms.
4 Approximated Attention
In addition to heuristic approaches aimed at restricting full attention computation, some research explores the mathematical essence behind attention kernel computations. These studies use estimation methods based on the sparsity or low-rank properties of attention matrices to approximate attention with linear complexity, albeit at the cost of precision. We introduce several of these approximation techniques below.
Low-Rank Approximation. Linformer employs Singular Value Decomposition (SVD) to approximate the attention matrix with a low-rank matrix , leveraging its low-rank property as proved in its Theorem 1. This approach involves two learnable projection matrices and of dimensions , where . The process includes projecting using respectively, followed by standard MHA kernel on with the projected . According to its Theorem 2, this low-rank technique approximates full attention with linear complexity while allowing for an error of .
Sparse-Kernelized Hybrid. Furthermore, inspired by Robust-PCA , the recent Scatterbrain provides a more accurate yet efficient approximation by combining LSH-based sparse matrices like Reformer’s and low-rank kernelized decomposition with randomized feature maps like Performer’s , as simplified in Eq.15, where we omitted the normalization step and the causal mask applying function. Not only does this method unify two approximation techniques to achieve linear time complexity of with higher precision, but it also offers flexibility to leverage various low-rank and sparse approximation methods as sub-components, more than just the example combination of Reformer and Performer.
5 IO-Aware Attention
All of the methods above in pursuit of efficient attention can be considered as trading off high attention quality for low computation complexity, based on some theoretical or empirical properties of attention matrix and NLP tasks, including locality, sparsity, low-rankness, and other heuristic or mathematical tricks. In comparison, these IO-aware attention mechanisms below collectively represent efforts to optimize attention computations by considering the memory bottleneck while preserving the exactness of attention kernel calculations.
Memory-Efficient Attention. This simple method is firstly proposed in , which utilizes the lazy softmax algorithm and tracks normalization factor to compute standard and numerically stable attention by sequentially processing each single/chunked-query attention. Such a simple method only needs constant working memory with respect to sequence length, while the time complexity is still quadratic.
SCFA. Although Flash Attention can be easily extended to support block-sparse structures , it may lack flexibility for handling other sparse strategies with irregular structures and arbitrary attention masks. The effort of SCFA extends the Flash Attention GPU kernel to accommodate a broad range of attention sparsity patterns, including key/query dropping and hashing-based attention like Reformer . This extension leads to a training speedup of 2.0 to 3.3 times without sacrificing perplexity, according to the report from the paper.
Paged Attention. While Flash Attention has effectively tackled the training memory bottleneck, LLMs still face challenges related to the memory consumption of the KV cache during inference, which grows dynamically with batched requests. Recognizing the memory wastage due to fragmentation and redundancy, vLLM proposes Paged Attention . This technique efficiently manages KV cache memory to minimize waste and allows flexible sharing across batched requests, drawing inspiration from memory paging techniques in virtual memory operating systems .
Long-Term Memory
Transformer architectures often struggle with capturing long-term dependencies due to in-context working memory, as highlighted in Sec.2.2. Researchers have explored two main avenues to address this challenge without compromising the advantages of full attention. First, inspired by RNNs, some have introduced recurrent mechanisms into attention by incorporating internal memory caches accessible through attention layers. This approach enables the model to maintain and retrieve information over longer sequences, compensating for the inherent lack of built-in long-term memory. Second, an alternative approach involves leveraging existing models as interfaces to external knowledge bases, such as specific documents or datasets. During inference, the model can read from these knowledge bases to enrich its contextual input and write to them from the user’s response to refresh its long-term memory. By integrating external knowledge in this manner, the model gains access to a broader range of context, enhancing its ability to handle long-term dependencies effectively.
Recalling the temporality of natural language representations instead of the success of full parallelism in Transformer, we introduce the concept of Internal MemoryCache based on recurrence mechanisms. It divides long text into a stream of fixed-length segments and enhances the query of the current -th segment in the -th layer with more contextual information . This contextual information is obtained from cached or distilled information from previous segments, stored in a memory cache denoted as , as shown in Eq.17. To facilitate later explanations, we assume that each segment has the same length , and the models consist of layers of transformer blocks. The notation represents the concatenation operation along the length dimension. It’s worth noting that the variables in the memory cache are usually detached from the computation graph, eliminating the need for gradient computation, which we denote with a hat accent, such as .
Segment-Level Recurrence. The segment-level recurrence is initially introduced into Transformer from Transformer-XL . As illustrated in Eq.18, it caches the output of previous consecutive segments in the last layer and concatenates them into the current segment in the present layer to extend the context for the current query. Such mechanism allows for extending the largest possible dependency distance to , where can be set as far as GPU memory allows. Building upon Transformer-XL, Segatron introduces the segment-aware mechanism by enhancing the token-level PEs combined with sentence-level and even paragraph-level ones. To further extend the dependency with multi-grained memory caching, Compressive Transformer stores the first FIFO fine-grained memory queue for previous segments as Transformer-XL does. However, instead of discarding old memory, it applies a compression function with the rate to compress it along the length dimension and pushes it into a secondary FIFO coarse-grained compressive memory queue of size . Combining these two types of memories, one can obtain longest context dependency as , as shown in Eq.19.
Retrospective Recurrence. Note that both Transformer-XL and Compressive Transformer deploy a shifting-one-layer-downwards recurrence by default, thus the maximum effective context length is limited by . To address it, similar to Feedback Transformer , ERNIE-Doc proposes an enhanced recurrence mechanism, a drop-in replacement by concatenating the output hidden states of previous segments in the same layer, instead of the last layer, simply formalized as Eq.21. In this manner, no only the maximum effective context length can be implicitly expanded, but the past higher-level representations can be exploited to enrich future lower-level representations as well. Additionally, it employs a retrospective feed mechanism by feeding the segments twice, where the first time only skims each segment while the second one retrospects to enable bi-directional information flow, which resembles the mechanism in READTWICE .
Alternate Cache Designs. Except following the prepending style of Mem as the memory cache, RMT formalizes the memory cache as special [mem] tokens, prepended both at the start and the end of each segment, as shown in Eq.22. After processing each segment, the read/write tokens will be split from the output embeddings and the write tokens will be taken as the [mem] tokens for next segment. By leveraging such recurrence mechanism with global memory tokens, RMT is demonstrated to scale effective context size to 1M tokens . And Memorizing Transformer applies (key,value) memory cache only for the top attention layer, but with a large cache size without compression. Besides, instead of a simple FIFO cache to read memory, they use kNN algorithm to retrieve top-k most similar (key,value) pairs for each query to prepend to the local ones, as Eq.23 indicates. In contrast, Memformer reads and writes the memory cache fully leveraging variants of self-attention with a forgetting mechanism, to retrieve and retain the most significant information through long-range timesteps.
2 External MemoryBank
The previously discussed mechanisms enhance the vanilla stateless Transformer model with sequential recurrence by prepending extra hidden states of previous inputs from an internal memory cache. However, these mechanisms have certain drawbacks. Firstly, a slight change in the memory mechanism may necessitate retraining the model from scratch, not fully utilizing pre-trained LLMs that already possess a good representation of dependency across the context window, albeit not long enough. Secondly, as highlighted in Memorizing Transformer , they often encounter the problem of memory staleness, where older hidden states in memory cache may exhibit distributional shifts from the latest ones during training, limiting the effectiveness of memory augmentation.
As a solution, another retrieval-augmented mechanisms decouples the model itself from its long-term memory storage. They leverage the part before the language head as a well-performing contextual information encoder to store long sequences as an external memory bank in the form of embeddings. And during queries, the model retrieves information from this memory bank based on certain criteria and concatenates it to comprise in-context working memory in real-time.
Cosine-Based Retrieval Criteria. LangChain , an hot open-source framework for developing applications like chatbots, accepts user-specified local documentation in common readable formats, then vectorizes this documentation using off-the-shelf LLMs into a memory bank. During each user interaction, it retrieves the top-relevant contexts based on dot-product cosine similarity between user prompt embeddings and the stored contexts. It then prepends these external contexts to the prompt, providing the LLMs with a more related input to generate responses. LangChain offers an effective pipeline to harness the capabilities of off-the-shelf LLMs and enhance their long-term memory with a cost-effective, flexible, and dynamic mechanism. Several similar works have also designed external memory banks for QA-like tasks or chatbot applications.
Heuristic Retrieval Criteria. Except for leveraging the cosine similarity, RETRO retrieves from the BERT-embedded KV memory bank through kNN search based on distance. Also based on kNN search, Unlimiformer designs a general approach for any existing pretrained encoder-decoder transformer to index unlimited input sequences for the decoder to retrieve its top- keys to apply cross-attention. In contrast, SiliconFriend proposes an enhanced Memory Bank mechanism to keep tracking the long-chat history with the user and provide specialized responses, consisting of dialogue logging with timestamps, events distillation into a high-level summary, awareness of user’s personality portrait and memory refreshment with a rate of forgetting. Similar to this text-based memory, RecurrentGPT enables recurrent prompting and defines the recurrent computation graph with ChatGPT by simulating the Long Short-Term Memory mechanism in LSTM . Moreover, RecallM organizes and updates the memory as a dynamic concept-aware knowledge graph to improve more complex continual learning and temporal reasoning of the knowledge during chat. Inspired by the Davidsonian semantics , Ret-LLM stores and retrieves knowledge from the memory bank in the form of triplets like which means ”A and B have a relationship of R”, as a general read-write memory unit and employs fine-tuned Alpaca to treat memory read/write as text-based API calls.
Learnable Retrieval Criteria. Despite these heuristic designs, REALM pre-trains a latent neural knowledge retriever using MLM as the learning signal, that takes charge of retrieving knowledge from the large textual corpus. LongMem trains another Transformer-based SideNet to decouple the memory retrieval and fusion process from the pre-trained LLMs which are only responsible for encoding the (key, value) pairs into the memory bank.
Extrapolative PEs
Recognizing the need to push the inference length boundary beyond , the research community has made significant efforts in this direction. Notably, according to , they have determined that distractors are the primary cause of failures in length generalization in the case of parity task. These issues, however, can be mitigated considerably through approaches such as scratchpad prompting . Nevertheless, in this section, our focus remains on the undeniable role that current PEs play in length generalization in more general scenarios.
Before entering into the concrete approaches, we would love to provide some insights below to enhance the understanding of this minute but essential design in Transformer for sequential modeling tasks.
Rethinking PEs as -Encoding. Su revisits the sine and cosine basis functions of Sinusoidal PE and RoPE, considering them as approximated terms for the -encoding system to represent any position number , as shown in Eq.24. This approach employs fixed -bits, where represents the power basis of the wavelength or period of the trigonometric basis functions, which increases as a geometric series with the dimension goes deeper.
To gain a deeper understanding of this concept, we can draw a comparison between Eq. 24 and Eq.5,6. It becomes evident that the -th -bit of the representation of involves the division of the -th power of , followed by some sort of periodical operations ( in Eq. 24 and in Eq.5,6).
Length Extrapolation Dilemma. Prior to the era of Transformers, RNN-based language models were trained on shorter sequences but were expected to generalize effectively to longer contexts, a phenomenon referred to as length extrapolation or length generalization . Unfortunately, recent studies have highlighted a significant shortcoming of length extrapolation ability for Transformer-based language models. This causes the insufficient context length limit during inference when applying to real-world applications, as analysed in Sec.2.2.
In the original Transformer paper , there is few discussion regarding the design insights or theoretical interpretation of their Sinusoidal PE. This has led many researchers to question its necessity and effectiveness, especially the blame on the extrapolation deficit, which points to the same trigonometry-based RoPE as well. To understand the bad extrapolation caused by current trigonometric PEs, we investigate and summarize two insights from distint views as below:
From a mathematical view, as Su explains in his blog, extrapolation, which involves inferring the whole from local information, depends on the high-order smoothness of the function. However, to accommodate sufficient positional information, these PEs are designed as combinations of high-frequency oscillatory trigonometric basis functions. This choice makes it challenging for the models to generalize without specific learning during training stages.
From a training view, due to the wavelength or period of the basis functions increases exponentially, proportional to , training samples constrained by currently supported are typically too short for the rear low-frequency dimensions to span the entire periodic cycle. This suggests only a few dimensions perceive complete periodic information thus receiving sufficient training for extrapolation, and the boundary is defined as critical dimension in (e.g. for Llama2-4k , the critical dimension is only 92). Consequently, direct extrapolation becomes prone to failure when relying on these poor-learned low-frequency components.
2 Attention Bias
As alternative mechanisms to explicitly encoding positional information, attention bias have been explored to capture the sequentiality and temporality of natural language incorporated into attention kernel. As shown in Eq.25, the attention bias is depicted as a matrix, denoted as , which is added to the unnormalized attention weights matrix before applying the softmax operation. Each element of this matrix, indexed by , carries positional information encoded by a function . Thus, it is reasonable to regard the attention bias as a form of relative PEs.
Early approaches like T5 employ learnable attention bias, denoted as , which is independent for each head in each attention layer. However, they did not explicitly address the problem of length extrapolation. The breakthrough in recognizing and addressing the extrapolation problem comes with ALiBi . ALiBi introduces a negative causal attention bias heuristically, as shown in Eq.26, where is a head-specific slope fixed before training and decreases geometrically with the head index . ALiBi successfully maintains low perplexity levels when extrapolating inference tokens beyond up to 16.
Following the success of ALiBi, several variants emerged in the quest to improve extrapolative PEs for Transformer-based LLMs. KERPLE extended the ALiBi-style attention bias by considering it as a composition triangle kernel to self-attention. Two extra learnable scalar parameters were introduced to generalize the bias kernel, as shown in Eq.27. The authors of Sandwich reused the Sinusoidal PEs to form the attention bias in a RoPE-style, as illustrated in Eq.28, with as a hyper-parameter to tune. Interestingly, another method discussed by Su in his blog utilizes a super-baseline approach during inference, as illustrated in Eq.29. This method relies on a local causal attention mask, where each query attends to keys whose distances have not exceeded while still applying RoPE. According to Su’s experiments, this approach proves to be simple, low-cost and performs sufficiently well compared to the more elaborate designs mentioned earlier, thus referred as a super-baseline.
3 Extended RoPE
RoPE, as introduced in Sec.2.1, is a widely-used positional encoding scheme utilized in popular LLMs such as Llama, GLM, PaLM. It offers advantages such as relative distance decay, training stability, compatibility with linear attention, and better length extrapolation capabilities compared to the traditional Sinusoidal PE, as demonstrated in various experiments , albeit not that satisfactory. Therefore, several research works have aimed to extend RoPE using various strategies to enhance its length extrapolation capabilities.
Scaling Strategies. Recent approaches by simply scaling it to extrapolate the inference context length with minimal or no fine-tuning , has gained prominence within the community. In LEX , the authors introduce an extended causal RoPE called XPOS, which incorporates an additional exponential decay term to scale RoPE, as illustrated in Eq.30, where is a scalar hyper-parameter. Similar techniques have been previously employed in PermuteFormer to adapt to the linear attention introduced by Performer . Positional Interpolation (PI) applies linear scaling on each position number from to , densifying the representation space to extend the farthest length boundary by times (see Eq.31). This strategy is experimentally shown to be more stable and require fewer fine-tuning steps than direct extrapolation.
However, it’s evident that simple linear scaling may hinder the network’s ability to distinguish the order and positions of closely spaced tokens, as their distances are compressed by a ratio of . Borrowing from the Neural Tangent Kernel theory (NTK) , which suggests that deep neural networks struggle to learn high-frequency information when the input dimension is low and corresponding embeddings lack high-frequency components, NTK-aware Scaling RoPE (NTK-RoPE) combines high-frequency extrapolation and low-frequency interpolation. It scales using a coefficient to achieve equivalence when applying interpolation by a ratio of for the lowest frequency term, while keeping the scale for high-frequency terms (see Eq.32). Surprisingly, this nonlinear scaling can be directly applied to LLMs pre-trained with RoPE, such as Llama, without any further fine-tuning to extend the context length boundary. This approach has already been deployed in open-source LLMs like CodeLlama .
Inspired by NTK-RoPE, several enhanced scaling methods have emerged. To avoid performance degradation when is still within the ”max length,” Dynamic-NTK delays the application of the scaling trick until exceeds the current supported context length. It gradually increases the ratio dynamically as increases. This idea has also been implemented in models like Qwen-7B and the latest version of Llama2 . Another approach, NTK-mix RoPE , introduces multiple coefficients for to interpolate less as the frequency increases. In contrast, NTK-by-parts opts not to interpolate the higher frequency dimensions at all while always interpolating the lower ones. The authors of NTK-RoPE also propose YaRN , which combines NTK-by-parts with a ”length scaling” trick that scales and by a constant temperature factor . This method claims to outperform all previous methods based on NTK-RoPE, whether fine-tuned or not. Additionally, Giraffe introduces another scaling strategy known as ”Power Scaling” (see Eq.33), where the exponent controls the decay ratio of low frequencies. This approach ensures that high-frequency elements of the basis are less affected than poorly learned low-frequency elements.
Inspired by NTK-RoPE, several enhanced scaling methods have emerged. To avoid performance degradation when is still within the , Dynamic-NTK delays the applying of the NTK scaling trick until exceeds the current supported context length. And it gradually increases the ratio dynamically as increases. This idea has also been implemented in models like Qwen-7B and the latest version of Llama2 . And to generalize scaling for each dimension, NTK-mix RoPE introduces multiple coefficients for to interpolate less as the frequency increases. In contrast, NTK-by-parts opts not to interpolate the higher frequency dimensions at all while always interpolating the lower ones. Recently the authors of NTK-RoPE further proposes YaRN , which combines NTK-by-parts with a length scaling trick that scales by a constant temperature factor . This method claims to outperform all previous methods based on NTK-RoPE, whether fine-tuned or not. Additionally, Giraffe introduces another scaling strategy, namely Power Scaling, as shown in Eq.33, where the exponent controls the decay ratio of low frequencies.This approach ensures that high-frequency elements of the basis are less affected than poorly learned low-frequency elements. Despite these manual designs, CLEX also exploits a neural ordinary differential equation (ODE) to learn a continuous scaling series as a continuous dynamical system.
However, both Leaky ReRoPE and ReRoPE involve connecting two stages of scaling, and their gap cannot be bridged by any linear transformation. Therefore, they require two attention matrix computations for each stage and utilize a boolean matrix to stitch them together, significantly increasing inference cost and limiting the real length boundary. Worse still, they are currently not compatible with the Flash Attention to mitigate the high computational cost. To adapt ReRoPE with Flash Attention, we have re-implemented the Flash Attention forward kernel to incorporate ReRoPE based on the Triton framework , somewhat alleviating its high computational costThe experimental implementation can be accessed at the following URL: https://github.com/Strivin0311/long-llms-learning/blob/main/notebooks/flash_rerope.ipynb.. Additionally, Giraffe introduces another truncation strategy, namely Basis Truncation, as illustrated in Eq.35, where are cutoff thresholds. This approach preserves the high-frequency components of the basis while cutting off the low-frequency elements to near-zero constant values () and even zeros when low enough, resulting in less complex extrapolation for the low frequencies.
Rearrangement Strategies. Based on the insights discussed in Sec.5.1, it is clear that the PEs of rear positions are updated fewer times than the front position embeddings. Such imbalance may lead to improperly trained rear positions. Recently, this issue has been addressed to some extent by some simple but effective works. SHAPE randomly shifts absolute positions during training to achieve shift invariance. Random Padding involves moving a random number of padding tokens to the front of the input sequence during fine-tuning, balancing the updating times across all positions. Additionally, Randomized PE randomly sub-samples an ordered set of positions from a much larger range of positions than the sequence length () during training, allowing the model to handle a broader range of PEs more robustly. PoSE fine-tunes the model to adapt all relative positions of the target context window (e.g. 128k) by adding a distinct skipping bias term to the position indices of training samples (e.g. 2k) to simulate longer inputs.
In summary, research on extrapolative PEs is a promising and rapidly-developing field, aiming to enhance the LLMs’ ability to infer long contexts in real-world scenarios with an available setting during training.
Context Processing
Many of the methodologies discussed earlier propose intricate designs around the attention module in Transformer architecture, including efficient attention kernels (Sec.3), long-term memory mechanisms (Sec.4), and extrapolative PEs (Sec.5). In contrast, there exist simpler and more straightforward approaches that view pre-trained LLMs as black-box or gray-box models. These methods tackle the challenge of handling long-context inputs that exceed the model’s length limitation by making multiple calls to the model, ensuring that the actual input provided to the LLM in each call does not exceed . While these methods do not explicitly enhance the LLMs’ inherent ability to process long-contexts, they leverage the LLMs’ remarkable in-context learning capabilities to address the issue, albeit at the cost of increased computation and potentially less precise answers.
These methods primarily involve partitioning lengthy texts into multiple segments, followed by the selection of specific segments based on a predefined strategy. The goal is to ensure that the chosen segments can comfortably fit within the context window of LLMs, all while minimizing the loss of pertinent information from the original lengthy text that is relevant to the posed query. These approaches differ in two key aspects. Firstly, they diverge in how they define the selection criteria, which are used to assign a priority score to each segment. Secondly, they vary in their selection strategies, with some methods involving the simultaneous sorting of all segments based on their scores, while others employ an iterative, greedy selection approach, considering segments one by one.
LangChain employs three strategies when dealing with retrieved context that exceeds the maximum context length of LLMs. One of these strategies is referred to as Map Rerank, in which LLMs are required to independently output answers for each segment, along with a confidence score. The answer with the highest confidence score is selected as the final output. For a more nuanced approach, CogLTX introduces a multi-step reasoning mechanism known as MemRecall. During each reasoning step, two models are sequentially employed to score the context segments in a coarse-to-fine manner. The top-k segments with the highest scores are added to the final candidate queue, while the remaining segments are deferred to the next reasoning step until the candidate queue is filled. In contrast, LoBART utilizes the ROUGE-2 score to select the top-k contexts during training. For inference, it trains an additional Hierarchical RNN model to generate surrogate priority scores for context selection.
2 Context Aggregation
In contrast to the selection-based methods, approaches of this nature take into account the contributions of all context segments to the final answer, rather than selecting just one. Initially, these methods extract relevant information from each segment individually. Subsequently, they employ various fusion strategies to aggregate the retrieved information, ultimately arriving at the final answer. Notably, these approaches exhibit variations in two key aspects. The first pertains to the manner in which information is extracted from each segment, while the second revolves around the diverse fusion strategies employed to integrate information across all segments.
In the context of Encoder-Decoder architecture LLMs like T5 and BART, there exists a category of methods known as Fusion-in-Decoder (FiD) . These methods leverage both the encoder to extract information in the form of embedded hidden states and the decoder to attend to all contextualized representations in order to generate the final output. To illustrate, consider the example of SLED . In SLED, the process of creating a contextualized representation involves overlapping a small portion of each segment with the neighboring segments, effectively forming what can be termed as context paddings. Subsequently, each segment is independently encoded through the Encoder, and the resulting embeddings are concatenated to generate embeddings for the entire extended document, with the exclusion of the context paddings. Finally, the Decoder integrates these locally contextualized embeddings through cross-attention, achieving a coherent fusion of information.
For decoder-only LLMs, LangChain introduces two additional aggregation techniques, in addition to the selection strategy known as Map ReRank. The first of these is referred to as Map Reduce, which involves the simultaneous processing of each segment to obtain answers in parallel. These answers are then passed on to another LLM, which synthesizes them into a final summary. In contrast, the second approach, named Refine, operates by progressively refining answers throughout the processing of each segment. In this strategy, the answers obtained from the previous segments are cascaded with the current segment, serving as the prompt for further refinement. This iterative refinement process continues until the final segment is processed.
In addition to LangChain, another recent method, PCW , employs a similar approach for handling long-context inputs. PCW partitions the extended context into multiple smaller context windows, each with a maximum length denoted as , with representing the total length of the context and being the length of the task-related tokens in the query, along with the maximum number of new tokens to be generated. Within each context window, tokens attend to each other in parallel, with their position indices isolated within the range of . Subsequently, the task-related tokens, which include the last context position indices within the range , attend to all the context tokens. This process aggregates the parallel information from each context window and fuses it to generate the final answer.
Similarly, another approach proposed by Su is NBCE , which treats the parallel context windows as a series of independent conditions denoted as , aiming to approximate the logarithmic posterior probability for the tokens to be generated. To achieve this, Su leverages the classic Naive Bayes algorithm to simplify the formulation. As deduced in Eq.36, represents the likelihood conditioned on the -th context window, represents the prior, and const is a constant that solely depends on . The computation of this formula is straightforward using LLMs, provided that access to their logits is available. Furthermore, Su extends this formulation to a more general case, as expressed in Eq.37, introducing a hyper-parameter and a pooling operation denoted as pool, which can be either average or max. In this extended form, the original Naive Bayes formula can be considered as a specific instance where is set to and pool is defined as average.
NBCE and PCW can be seamlessly applied to any readily available open-access LLMs, allowing for significant extensions of the context length. However, it’s important to note that both methods operate under the assumption that the relationships among these context windows are negligible and can be treated uniformly in an unordered fashion. Consequently, their performance may suffer when confronted with tightly interconnected context windows exhibiting sequential relations or when dealing with an excessive number of windows to process in parallel.
Miscellaneous
This section offers a concise overview of miscellaneous solutions that extend the previous four categories discussed, providing a broader perspective on advancing effective context window of LLMs or the efficiency when utilizing off-the-shelf LLMs. It is important to note that the literature presented here may not be exhaustive or tailored for Transformer-based models. Indeed, many of these techniques are applicable universally to any model equipped with deep neural networks, though they are particularly crucial for large-scale LLMs. Therefore, we hope this section serves as a valuable guide to these diverse directions.
Specific Objectives. In contrast to the conventional pre-training objectives such as MLM or CLM discussed in Sec.2.1, recent research explores tailored approaches to adapt pre-training for specific tasks, aiming to enhance LLMs’ efficacy in capturing intricate long-range dependencies and discourse structures in longer texts compared to shorter ones . These approaches involve alternative loss functions, additional terms, and specialized pre-processing techniques. For instance, XLNet introduces a permutation objective that excels in various NLP tasks. ERNIE-Doc extends this approach to long documents with a Segment-Reordering Objective to model long-range relationships. DANCE employs a divide-and-conquer pre-processing strategy for summarization tasks, breaking the long document and its summary into multiple source-target pairs. PEGASUS introduces the Gap Sentence Generation (GSG) objective for abstractive summarization, while PRIMERA extends it across multi-documents using the Entity Pyramid method.
Mixture of Experts. The concept of Mixture of Experts (MoE) stands as a powerful augmentation for giant LLMs, by substituting the dense FFN layer with a MoE layer, which incorporates multiple specialized experts. Each expert excels in handling specific input types or tasks, and a dynamic gating mechanism arranges the selection of the most suitable expert for a given input. This approach can be implemented in various ways, including employing expert modules optimized for specific tasks , utilizing sparse activation and sharding to multiple devices , and adapting mixture weights through training to determine the contribution of each expert, as seen in Soft MoE . Then, routing mechanisms are responsible for selecting the top-k experts for each token based on their gate values, where determines the number of experts . However, in Switch Transformer , they have found that setting , referred as Switch Routing, can preserve model quality while reducing routing computation. Finally, the output is obtained through a weighted summation of the contributions from these selected experts. MoE techniques can significantly enhance versatility, reduce computational demands, and elevate the efficiency and effectiveness of modeling large-scale contexts.
Parallelism. Leveraging modern aggregated GPU memory within and across nodes, recent research has introduced various parallelism strategies to scale up model sizes and extend sequence length. We summarize commonly-used parallelism paradigms with brief introductions as follows:
Data Parallelism , widely integrated into PyTorch , is the most commonly-used way to accelerate training in a distributed manner across multiple devices. It replicates the model on each device to generate gradients independently and communicates them at each iteration to maintain consistency.
Tensor Parallelism introduces tensor splitting, where individual layers of the model are horizontally partitioned over multiple devices.
Pipeline Parallelism splits the model layers vertically along the batch dimension into different partitions of micro-batches on separate devices. Each device processes one micro-batch received from the previous one in a pipeline fashion.
Sequence Parallelism divides the input sequence into multiple chunks and feeds each chunk into its corresponding device. It incorporates ring-style communication for computing the attention output.
Expert Parallelism , as discussed earlier in MoE, places different experts on different GPUs and executes them in parallel. Classic all-to-all communication primitives are often used to implement this form of parallelism .
Note that, the combination of a),b),c) is often referred to as 3D parallelism , which efficiently scales models to trillions of parameters. However, 3D parallelism is primarily designed for larger model sizes rather than longer sequences. To train large-scale models with long sequences, integrating a),b),c) and d) into 4D parallelism is a promising approach. Additionally, memory optimization strategies like ZeRO can further enhance efficiency by reducing redundancies in conventional data/tensor parallelism.
Alternate Memory-Efficient Inference. Other approaches have been developed to enhance memory efficiency in large-scale Transformer-based LLMs during inference. These methods encompass weight pruning , weight factorization , weight quantization , weight partitioning and knowledge distillation . Among these, quantization strategies are particularly vital for practical deployment of massive LLMs, as they reduce parameter precision to alleviate memory demands and accelerate inference. Additionally, there are simpler strategies to alleviate the large KV cache during inference like Multi-Query Attention (MQA) and Grouped-Query Attention (GQA) . In particular, they reduce the number of heads for keys and values, and share them equally across multiple query heads. These approaches have already been applied into state-of-the-art LLMs such as GLM and PaLM, to the best of our knowledge.
Evaluation Necessity & Optimization Toolkit
In this section, we have explored the evaluation necessities commonly used for assessing the long-context capabilities of LLMs, which include datasets, metrics, and baseline models. Moreover, we have also investigated several popular optimization toolkits, such as libraries, systems, and compilers, aimed at enhancing the efficiency and effectiveness of LLMs throughout their development stages. For a concise and clear presentation, we have organized detailed information in corresponding tables located in the appendix A, B, C, D. Here, we provide an overview of these tables and their content.
Datasets. This part presents a summary of the evaluation datasets used to assess the long-context capabilities of LLMs. Detailed dataset information, including language, task types, length statistics, quality, splits, file size, sample count, and format, is provided in the Tab.1 located in Appendix A. We illustrate the meta information about the table as follows:
Language: Language information for each dataset is represented using abbreviations, such as en for English, zh for Chinese, and code for programming languages. Multiple languages within a dataset are concatenated using ””.
Task Amount and Types: This property defines each dataset’s tasks and divides them into 10 categories, including language modeling (LM), multi-choice question-answering (MCQA), extractive question-answering (ExtQA), summarization (Summ), text classification (Class), reasoning tasks (Math), and natural language generation (Cloze, GenCloze is to predict and fill the short blanks in some long contexts like code filling, and Gen is more about the tasks to open writing like story generation, whose context may not be that long but the generated part is usually quite extensive.).
Length: The average (avg), minimum (min), and maximum (max) sample lengths are provided in kilo ”words”In the context of our study, ”words” are approximately considered to be separated by spaces in English and code, while individual Chinese characters are treated as words. for each dataset, where ”words” are defined based on sample contentFor example, if one typical sample has the prompt template like ”Read this context, and answer the question below: question”, we will calculate the number of words in both context and question part, ignoring the fixed remaining part in the template..
Quality: Quality assessment is simply based on two sub-dimensions: Human Labeled (labels generated by humans) and Model Gen (prompts or labels generated by LLMs).
Splits: This indicates how datasets are partitioned, including conventional triple-split formats like train/test/val, a single test split only for evaluation purposes such as LongBench .
Size and Count: These two provide statistics on file size (in bytes, B) and sample count for each split.
Format: It tags the file format of samples, including common formats like jsonl, json, csv, txt, and others.
Metrics. This section provides a concise summary of 9 categories of general evaluation metrics commonly used across the 10 NLP task types categorized earlier. The detailed metrics for each task are presented in a Tab.2 located in Appendix B. Below, we offer a brief introduction to these metrics.
CE/PPL (Cross-Entropy/Perplexity): CE is the most commonly-used loss function for pretrained language models, which quantifies the divergence between predicted distributions and the true distribution approximated from the large training text corpus. And PPL measures how well a language model predicts a sequence, simply formalized as , where denotes the cross-entropy loss for the test set.
Acc/F1 (Accuracy/F1 Score): Accuracy measures the proportion of correct predictions in tasks with objective answers like classification, MCQA, while the F1 Score balances precision and recall for imbalanced datasets.
EM (Exact Matching): It evaluates whether a generated sequence matches the reference exactly, used in tasks that require absolute precision like code auto-completion.
ROUGE-1/-2/-L : ROUGE scores assess the similarity between text pairs by comparing n-grams overlapping, typically setting n=1 (unigram), 2 (bigram), and L (longest). They are widely used in tasks EM may fail, such as summarization.
BLEU/METEOR/TER : These metrics are specific in machine translation tasks. BLEU measures the overlap of generated and reference translations based on n-grams. METEOR evaluates translation quality by considering various linguistic factors. TER quantifies the edit distance between the generated and reference translations.
EntMent (Entity Mention) : It evaluates the coverage and correctness of important entities, such as people, organizations, locations, etc, mentioned in the generated text, like the summary in the summarization task.
Pass@k: This metric evaluates whether a generated answer ranks within the top-k answers provided by a model, commonly used in code generation and some math tasks with multiple possible solutions.
Human/Model Judge: This involves human or power models like GPT-4 to score text quality based on fluency, coherence, and other subjective criteria suitable for tasks like story generation.
Baselines. In this part, we gather pretrained or fine-tuned LLMs frequently used in the literature to serve as baselines for assessing long-context capabilities in some downstream tasks. We provide an overview of these models about their basic information (name, base architecture, , parameter sizes), special features (training techs and fine-tuning task), and available links for publication (huggingface, GitHub, homepage, and paper) to Tab.3 in Appendix C to facilitate researchers’ convenience.
Toolkits In this section, we present a collection of valuable toolkits, encompassing libraries/packages, compilers, systems/frameworks, etc, designed to enhance the efficiency and effectiveness of LLMs throughout their development life-cycle. For a quick overview of these toolkits, please refer to Tab.4 provided in Appendix D.
Discussion
While significant progress has been made in this field as we discussed in sections 3, 4, 5, 6, several challenges still remain. Therefore, in this section, we delve into the these pivotal challenges and propose potential directions for future research and development in the context of enhancing long-context capabilities within Transformer-based LLMs, with a specific focus on architectural enhancements.
Attention Trade-off. As discussed in Sec.3, the efficient attention methods we have explored often involve a delicate trade-off between maintaining full-scale attention dependencies (such as local attention) or achieving higher attention score precision (e.g., through approximated attention) to mitigate the computational demands of standard attention kernels. However, with longer contexts, the discourse structure and interrelated information become increasingly complex, necessitating the ability to capture global, long-range dependencies while preserving precise relevancy. Addressing this challenge requires finding an optimal equilibrium between computational efficiency and retaining as much precision in attention patterns as possible. Thus it remains an ongoing pursuit in the field of long-context LLMs. Fortunately, some recent innovations, such as Flash Attention , explore IO-aware solutions beyond algorithmic level, which extremely improve efficiency in both runtime and memory overhead without any loss of attention precision. This is an exhilarating potential avenue to tackle this issue in practical applications. And furthermore, one can explore the integration of preceding efficient strategies with these drop-in replacements, leveraging powerful GPU kernel programming tools like Cuda, or the more light-weighted Triton , as exemplified in SCFA .
Memory Effectiveness and Efficiency. As discussed earlier in Sec.2.1, 2.2, we have outlined the limitations arising from the absence of an explicit memory mechanism, relying solely on in-context working memory, as well as the significant increase in KV cache memory consumption during extended context interactions. These challenges collectively underscore the need for more effective and efficient memory mechanisms within the realm of Transformer-based LLMs. While we introduced various long-term memory mechanisms in Sec.4, they are constrained by the additional memory overhead introduced by their intricate heuristic design, thus leading to potential performance degradation over time due to frequent refreshment. To tackle this challenge, researchers can investigate more efficient strategies for organizing memory storage and enhancing read/write throughput, drawing inspiration from recent advancements such as Paged Attention .
Length Extrapolation Mining. In Section 5, we conduct a thorough analysis of the challenges associated with length extrapolation in Transformer-based models, with a primary focus on the prevalent design of positional embeddings. And we offer a holistic overview of the very recent breakthroughs, particularly the extended strategies applied on the RoPE , which, as we believe, holds significant promise in addressing the extrapolation limitation. However, it’s important to note that these advancements often rely on simplified observations of complex high-dimensional positional embedding properties and incorporate straightforward heuristic adjustments. This prompts us to question the theoretical foundations of modeling sequentiality using high-dimensional embeddings and to explore the potential resurgence of learnable embeddings, guided by these heuristic designs with many hyper-parameters to tune. We believe that future research should delve deeper into this area, as exemplified in CLEX , especially in terms of developing a robust theoretical framework for modeling sequentiality under Transformer settings.
Specific yet Universal Objective. While we have discussed specific objectives tailored for long-text modeling, it is worth noting that many of them are limited to certain types of tasks or are only compatible with the MLM objective rather than the more common CLM objective nowadays. This highlights the need for specific yet universally applicable causal language modeling objectives that can effectively capture long-range dependencies from the early stages of model training. This could potentially be achieved by aligning such objectives with the mining of an effective PE scheme, as mentioned earlier.
Reliable Metric Demand. Speaking of evaluation metrics, we have investigated many options in Sec.8. However, drawing from our prior experience in evaluation, it is evident that commonly-used metrics such as ROUGE scores often exhibit significant disparities when compared to human judgment scores, which can be seen as the oracle. With the rapid deployment of LLMs in real-world scenarios, there is an increasingly pressing need for more dependable metrics to assess long-context capabilities, particularly in generative tasks where precise ground truth is elusive. One promising avenue involves leveraging the robustness of state-of-the-art LLMs like GPT4 to serve as substitutes for human judges, although the associated high costs continue to pose challenges for wider adoption within the research community.
By addressing these challenges and exploring future potential solutions, researchers can pave the way for more capable LLMs that excel in understanding and processing long-context information, opening up new opportunities across various real-world applications.
Conclusion
In this survey, we comprehensively navigate the landscape of architectural advancement in Transformer-based LLMs, to enhance the capabilities handling extensive context windows across various development stages, with a holistic taxonomy that categorizes the these methodologies targeting different module designs in Transformer. Then we explore evaluation necessities specific for long-text tasks and some optimization toolkits that integrate many tools to augment LLMs’ efficiency and efficacy. We further identify key challenges with corresponding future directions. In addition, our repository ensures that readers stay updated with the latest research in this dynamic field. As LLMs continue to evolve rapidly, we sincerely hope our survey serve as a valuable resource for researchers seeking to harness their power in building powerful long-context LLMs, ultimately advancing the pursuit of the era of AGI.
Acknowledgement
We would like to express our gratitude to Zenan Li and Hao Gao for their helpful discussions and feedback during the early stages of this paper. We also extend our appreciation to our Baidu colleagues, including Hao Chen, Linyun Liu, PengHao Zhao, and others, for their contributions and insights in the initial research phases.
Additionally, we acknowledge the generous support from the Baidu AI Cloud Group (ACG). We are especially thankful to Dou Shen, the Executive Vice President and Head of ACG, for the great idea and gracious invitation of the first session of Baidu ACG Summer Camp. This opportunity has been instrumental in shaping our research and providing valuable experiences.
References
Appendix A Datasets
Note that: the presence of common dirty data may result in extremely short samples, thus many datasets in the table containing samples with a minimum length approaching zero.
Appendix B Metrics
Note that: the ✗ in the table does not imply that a specific metric cannot be applied to a task. Rather, it suggests that the metric might be less commonly used or that there could be more suitable alternatives.