Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems

Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Hongyi Jin, Tianqi Chen, Zhihao Jia

Introduction

Generative large language models (LLMs) have become a driving force behind significant advancements in artificial intelligence (AI) and have demonstrated exceptional performance across a wide range of language-related tasks. From machine translation to sentiment analysis, question answering, and text generation, these models have shown their prowess in understanding, generating, and manipulating human languages. The advent of Transformer-based architectures, such as GPT-family (Generative Pre-trained Transformer) (OpenAI, 2023), LLaMA-family (Touvron et al., 2023), and other latest public LLMs (e.g., OPT (Zhang et al., 2022), BLOOM (Workshop et al., 2022), Mistral (Jiang et al., 2023a), DeciLM (Team, 2023), Baichuan (Yang et al., 2023c), GLM (Zeng et al., 2022)) has played a pivotal role in this paradigm shift, revolutionizing the way natural language processing (NLP) tasks are approached. Beyond NLP, these models are also transforming a wider range of applications, including automated programming (Chen et al., 2021b), science discovery (Jo, 2023), personalized digital assistants (Dong et al., 2023), creative arts (Ramesh et al., 2021), and next-generation computing architecture (Packer et al., 2023), demonstrating their versatility and profound impact across various industries.

However, the unprecedented success of LLMs has also given rise to several challenges, most notably, their formidable computational requirements during serving. The immense model size and complexity, coupled with the need for extensive computational resources, have impeded their widespread deployment in real-world applications. The resource-intensive nature of these models raises concerns over energy consumption, scalability, and accessibility, hindering their adoption in broader communities without rich compute resources like large companies.

This survey paper aims to address the critical need for efficient LLM serving and presents an exhaustive exploration of the existing multifaceted strategies proposed by the research community to tackle this challenge. We present an in-depth examination of the entire spectrum of solutions, spanning from algorithmic innovations to novel system architectures, all aimed at optimizing the inference process for large language models.

The primary objective of this survey is to provide a comprehensive overview of the latest advancements in LLM serving and inference. We will systematically review and categorize the existing techniques based on their underlying approaches, highlighting their strengths and limitations. The survey will cover a broad range of methodologies, encompassing decoding algorithm, architecture design, model compression, low-bit quantization, parallel computation, memory management, request scheduling, and kernel optimization.

2. Structure

The paper is structured as follows: Section 2 introduces the background information about LLM serving. Section 3 includes our taxonomy of existing approaches on efficient LLM serving and revisits these related works from two aspects: algorithmic innovations (§ 3.1) and system optimizations (§ 3.2). After that, we list some representative LLM serving frameworks and provide analysis in Section 4. Section 5 discusses benchmarks of LLM serving systems. Section 6 clarifies the connection between this survey and other related literature. Finally, we propose some promising exploration directions in Section 7 for improving generative LLM serving efficiency to motivate future research.

Background

Transformer-based Large Language Models (LLMs) have marked a significant shift in the field of natural language processing, introducing a new paradigm for understanding and generating human language. Central to this innovation is the Transformer architecture, which is built upon the concept of self-attention mechanisms (Vaswani et al., 2017), allowing the model to weigh the importance of different parts of the input data when making predictions. Mathematically, the self-attention mechanism in Transformers can be described as follows: For an input sequence X=[x1,x2,...,xn]X=[x_{1},x_{2},...,x_{n}], the Transformer computes a set of queries QQ, keys KK and values VV using linear transformations of XX. The self-attention scores are then computed as:

where dkd_{k} is the dimension of the keys. This mechanism allows the model to focus on different parts of the input sequence for each element of the output, capturing complex dependencies regardless of their distance in the input sequence.

Another important structure in Transformers is the Feed-Forward Network (FFN), which is present in each layer of the Transformer and significantly contributes to its computational intensity. The FFN typically consists of two linear transformations with a non-linear activation function in between, usually represented as:

Here, W1W_{1}, W2W_{2}, b1b_{1}, and b2b_{2} are learnable parameters of the FFN, and the non-linear function max⁡(0,⋅)\max(0,\cdot) (ReLU, in this case) introduces the necessary non-linearity into the model, allowing it to learn more complex patterns. The FFN is responsible for a significant portion of the model’s parameter count and, consequently, its memory footprint and computational load. In each Transformer layer, after the multi-head attention (MHA) aggregates information from different parts of the input, the FFN processes this aggregated information independently for each position. This parallel processing capability is a key strength of the Transformer, allowing it to handle sequences effectively. However, it also means that the computational load and memory requirements scale with the length of the input sequence and the depth of the network.

The combination of self-attention and FFN in Transformer-based LLMs enables these models to capture a wide range of linguistic contexts and nuances, setting new benchmarks in various NLP tasks. However, the substantial computational requirements for training and inference have become a critical area of research, focusing on optimizing these aspects without significantly compromising performance. The Transformer model also includes other key components like position encoding, which adds information about the position of each token in the sequence, and the multi-head attention mechanism, which allows the model to focus on different parts of the sequence in different representational spaces.

2. GPUs and Other Accelerators

The rapid advancement of LLMs owes much to the evolution of GPU architecture and other accelerators, which are integral to enhancing model performance and efficiency. GPUs (Graphics Processing Units) have emerged as a cornerstone in this field, primarily due to their superior parallel processing capabilities. Unlike traditional CPUs, which are designed for sequential processing, GPUs consist of thousands of small, efficient cores designed for handling multiple tasks simultaneously. This makes them exceptionally well-suited for the matrix and vector operations that are ubiquitous in deep learning computations, especially for Transformer-based models.

A typical GPU architecture comprises an array of Streaming Multiprocessors (SMs), each containing several cores that share a common instruction unit but can execute independent threads in parallel. Additionally, the shared memory (SRAM) within each SM allows for efficient data exchange and synchronization among threads, significantly optimizing the memory access patterns required in LLM computations. This design is particularly beneficial for the computationally intensive tasks in LLMs, such as the calculations of self-attention and feed-forward networks in Transformers. GPUs also come equipped with high-bandwidth memory (HBM), which allows for faster data transfer rates, significantly reducing the bottleneck associated with memory access during large-scale computations. Moreover, the latest GPU architectures, such as NVIDIA’s Ampere and Hopper architectures, continue to offer enhancements and push the boundaries of LLM computation, such as improved memory bandwidth and capacity, higher floating-point operations per second (FLOPS), specialized mixed-precision computing units (i.e., Tensor Core) and more efficient utilization of resources, further accelerating the performance of LLMs. Some of them support various precision formats, including FP32 (32-bit floating point), TF32 (TensorFloat-32), FP16 (16-bit floating point), BF16 (Brain Floating Point), and even INT8/INT4, allowing for flexible trade-offs between computational speed and numerical precision, essential in optimizing LLM performance.

Beyond GPUs, a vast array of hardware platforms have been explored for LLM deployment, encompassing CPUs (Shen et al., 2023; int, 2023), mobile and edge devices (Dettmers et al., 2023b), ASICs (Peng et al., 2023b; Zhou et al., 2022c), as well as specialized accelerators such as TPUs (Jouppi et al., 2023), FPGAs (Yemme and Garani, 2023), and other emerging AI chips from various manufacturers (e.g., Apple M2 Ultra (Lai et al., 2023), AWS Inferentia (aws, 2023), SambaNova (sam, 2023), Cerebras (Dey et al., 2023), Graphcore IPUs (gra, 2023)). This survey primarily underscores research anchored in the use of GPUs, and several technical motivations drive this emphasis. Due to their architectural innovations and superior computational power, GPUs have dominated the research area of large-scale deep learning in the past few years (Ben-Nun and Hoefler, 2019). Furthermore, the programming languages of GPUs, like NVIDIA’s CUDA and AMD’s ROCm, facilitate a fine-grained control over thread hierarchies, allowing researchers to exploit the massive parallelism inherent in GPUs. It attracts numerous developers to build mature software ecosystems on top of these GPUs, fostering a majority of the seminal and advanced LLM research. While other hardware platforms indeed bring unique strengths to specific contexts, the vast reservoir of research, development, and deployment centered around GPUs makes them an indispensable reference for an in-depth comprehension of LLM inference methodologies. Considering the hardware similarities, other hardware platforms can also benefit from the design philosophies, insights, and methodologies discussed in this survey.

3. LLM Inference

LLM inference, particularly in models like GPT (Generative Pre-trained Transformer), often employs an auto-regressive decoding approach. This method is central to how these models generate text, ensuring that each new word or token produced takes into account the entire sequence generated so far. Auto-regressive decoding operates under the principle of sequentially predicting the next token in a sequence, given all the previous ones, as shown in Algorithm 1.

Here, P(y∣Xt−1)P(y|X_{t-1}) represents the probability of the next token yy given the current sequence Xt−1X_{t-1}, and ⊕\oplus denotes the concatenation operation. The argmax function is used to select the most probable next token at each step.

This auto-regressive approach is fundamental in LLM inference for generating coherent and contextually appropriate text. It ensures that each token generated is conditioned on a comprehensive understanding of all previously generated content, allowing LLMs to produce highly relevant and fluent text sequences. Prior studies have provided in-depth analysis on the algorithmic intensity of Transformer-based LLM inference (e.g., counting the FLOPS, I/O and memory consumption) and extensive empirical results on cost estimation (e.g., modeling the inference latency (Chen, 2022)) according to the auto-regressive decoding algorithm execution. The optimization of LLM inference is a complex problem as there can be different optimal strategies with different algorithm configurations and system setups.

4. Challenges

This section describes a variety of challenges for efficient LLM serving.

Efficient large language model inference requires achieving low-latency and fast response times, especially in real-time applications like chatbots, virtual assistants, and interactive systems. Balancing model complexity with inference speed is a critical challenge that necessitates optimizing algorithms and system architectures to minimize response time without compromising accuracy.

Large language models come with significant memory requirements due to their size and the vast number of parameters they contain. Deploying such models on memory-constrained devices poses a challenge, demanding the development of effective model compression techniques and system optimizations to reduce memory footprint without sacrificing performance.

Inference systems often face varying levels of request loads in production environments. Ensuring scalability and high throughput to handle multiple simultaneous requests efficiently requires parallel computation, request scheduling, and other system-level optimizations to distribute computational workload effectively across resources.

Efficiently leveraging hardware resources is crucial for large language model inference. Adapting LLM models to diverse hardware platforms and architectures, including CPUs, GPUs, and specialized accelerators, demands hardware-aware algorithm design and optimization to exploit the full potential of the underlying hardware.

Optimizing the efficiency of LLM inference may sometimes involve trade-offs with model accuracy. Striking the right balance between model size, computational complexity, and performance is a challenging task that requires careful consideration and evaluation of various algorithmic and system-level techniques.

Taxonomy

Existing efforts on improving the LLM serving efficiency can be broadly classified into two categories, including algorithmic innovations and system optimizations, which will be discussed individually.

This section presents a comprehensive analysis of the various algorithms and techniques proposed to optimize language model inference efficiency. These works are proposed to address the native performance flaws of large-scale Transformer models through algorithmic advancements.

In this section, we review novel decoding algorithms as shown in Figure 2 that optimize the inference process of LLMs. These algorithms seek to reduce computational complexity and enhance the overall efficiency of language model inference during generation tasks.

Non-autoregressive decoding. A major limitation of existing LLMs is the default auto-regressive decoding mechanism, which sequentially generates output tokens one by one. To address this issue, one representative line of work is to abandon the autoregressive generation paradigm and decode the output tokens in parallel. Non-autoregressive decoding (Gu et al., 2018; Guo et al., 2019b; Ghazvininejad et al., 2019) is first proposed for machine translation acceleration by breaking the word dependencies during decoding and assuming a certain degree of conditional independence. To alleviate the translation quality reduction, some follow-up studies like semi-autoregressive decoding (Ghazvininejad et al., 2020), further extend these non-autoregressive methods to reach auto-regressive model quality by modeling output dependencies (Gu and Kong, 2021; Zhan et al., 2023) or iteratively refining output tokens (Lee et al., 2018). Blockwise parallel decoding (Stern et al., 2018) inserts a single feedforward layer to the base LLM to make predictions for multiple future positions in parallel, then backs off to the longest prefix validated by the base model. However, these approaches require to costly reproduce a new LLM with the new dependencies or tune partial layers of the original LLM, which are not always possible. Some recent efforts have been dedicated to generate multiple tokens at one decoding step without any training or modification to the model. Parallel decoding (Santilli et al., 2023) reframes the greedy auto-regressive decoding as a system of nonlinear equations solvable in parallel leveraging Jacobi and Gauss-Seidel fixed-point iteration methods for fast inference. A thorough survey on non-autoregressive translation (Xiao et al., 2023b) has been proposed to summarize the recent advances in this direction. Until now, due to the unawareness of the conditional dependence between output tokens, the output quality of most of non-autoregressive methods has been still less reliable than the auto-regressive method despite an improvement in decoding speed.

Speculative decoding. Another line of work addresses the sequential execution limitation by leveraging speculative execution (Burton, 1985) and improving decoding parallelism. Each decoding step during the autoregressive LLM inference process can be treated as the execution of a program with conditional branches, such as deciding which token to generate next. Speculative decoding (Leviathan et al., 2023; Chen et al., 2023a) has been proposed to make decoding predictions of multiple steps first in an efficient manner (e.g., using a smaller draft model with fewer model parameters) and verify these predictions simultaneously with the LLM. However, there are still several practical challenges remaining when applying speculative decoding to LLMs, e.g., how to make decoding predictions light-weight and accurate enough and how to achieve efficient parallel verification using LLMs. SpecInfer (Miao et al., 2023a) first addresses these challenges by introducing multiple small draft models coupled with a novel tree-based speculative inference and token verification mechanism (which are directly adopted by (Xu et al., 2023d; Cai et al., 2023; He et al., 2023; Spector and Re, 2023; Sun et al., 2023c; Monea et al., 2023; Liu et al., 2023b; Zhou et al., 2023)) and proposes a low-latency LLM serving system implementation (§ 4). The main advantage of speculative decoding is that it increases the parallelism without any changes to the outputs. Such guarantee comes from the fact that the predicted output is always verified by the original LLM and the fallback mechanism (Kim et al., 2023d) takes effect when prediction goes wrong.

Early exiting. Some other studies attempt to utilize the deep multi-layer architecture of existing LLMs and leverage the early exiting (Teerapittayanon et al., 2016) mechanism to accelerate the decoding process. The intuition is that the output of early model layers has the potential to infer the target distribution confidently. They can emit predictions based on internal classifiers instead of running the whole LLM, and various exit conditions have been explored (Xin et al., 2020; Liu et al., 2020; Zhou et al., 2020; Liao et al., 2021; Sun et al., 2022; He et al., 2021; Kong et al., 2022; Ye et al., 2021; Zeng et al., 2023). They are also called by adaptive computation (Schuster et al., 2021; Del Corro et al., 2023) since they adjust the amount of computation per request to amortize the total inference cost, i.e., taking less computation for easier inference requests. Broadly, these approaches are mostly restricted to the insufficient information carried by internal representations and may not faithfully making accurate predictions.

Cascade inference Driven by the varying complexities of inference requests, cascade inference employs a suite of LLMs of differing scales to minimize response time. Instead of directly using a massive model for every query, CascadeBERT (Li et al., 2020) involves a series of internal classifiers corresponding to different model depths, organizes them in a cascading manner and adaptively selects proper ones based on the instance difficulty. Tabi (Wang et al., 2023) optimizes for serving discriminative models (i.e., not generative LLMs), but it takes a similar approach to incorporate small models and LLMs to handle queries with different confidence. FrugalGPT (Chen et al., 2023d) leverages a learning-based approach to adaptively assign queries to different LLM APIs, optimizing both cost and performance. A concurrent work (Zhu et al., 2023c) jointly optimizes model multiplexing and query caching and also analyzes the optimality of minimizing inference cost. Mixture-of-thought (Yue et al., 2023) extends the cascade idea to LLM reasoning tasks for cost-saving, which samples answers from both Chain-of-Thought (Wei et al., 2022) and Program-of-Thought (Chen et al., 2022) prompts. Overall, cascade inference is a promising direction for enhanced inference efficiency, but it is still challenging to design an accurate dispatching mechanism to avoid compromising model quality.

1.2. Architecture Design

This subsection explores innovative architecture designs tailored for large language models. Researchers have proposed novel model architectures (He and Hofmann, 2023) beyond the original Transformer that strike a balance between model size, performance, and efficiency, opening new avenues for faster and resource-efficient inference.

Configuration downsizing: To reduce the computation cost of LLM inference, a straightforward approach is to downsize the model configurations, such as using shallow encoders (Goyal et al., 2020; Modarressi et al., 2022) or decoders (Kasai et al., 2020), weight sharing, and vocabulary shrinking (Shi and Knight, 2017). However, reducing the number of model parameters also affects the downstream tasks’ performance.

Attention simplification: One prominent challenge associated with self-attention calculations is the computational complexity O(L2)\mathcal{O}(L^{2}), which scales quadratically with the input sequence length LL. Numerous Transformer variants (Tay et al., 2023) have been proposed to simplify the standard attention into more efficient alternatives for very long sequence tasks, such as sparsification (Zaheer et al., 2020), kernelization (Katharopoulos et al., 2020), and factorization (Wang et al., 2020a). Recently, there is a trend of borrowing the ideas from prior attention simplification approaches, generalizing and combining them to shorten the context and reduce the size of KV cache, as well as the attention complexity, with slightly decoding quality degradation (e.g., sliding window attention (Jiang et al., 2023a; Zhang et al., 2023c), hash-based attention (Pagliardini et al., 2023), dilated attention (Ding et al., 2023)). One category of these approaches is context compression by compressing the context into fewer soft tokens (e.g., replacing with summary tokens (Chevalier et al., 2023) or landmark tokens (Mohtashami and Jaggi, 2023), leveraging additional autoencoder schemes (Liu et al., 2023c; Ge et al., 2023a)) or directly dropping or rephrasing unimportant context tokens based on different importance guidance (Li et al., 2023a; Jiang et al., 2023b; Mu et al., 2023; Fei et al., 2023) (or called semantic compression). For example, adaptively sparse attention (Anagnostidis et al., 2023) takes a learning-based approach to eliminate uninformative context tokens dynamically for each token. Scissorhands (Liu et al., 2023a) and H2O (Zhang et al., 2023d) select a few important tokens that might have a substantial influence for future decoding process and save their KV cache. StreamingLLM (Xiao et al., 2023a) values the initial tokens and maintains them with the sliding window, which is also similar to prior work (Beltagy et al., 2020). FastGen (Ge et al., 2023b) allows different attention heads to employ different emphasizing patterns adaptively. Table 1 illustrates the sparse attention patterns of four representative categories of approaches and their applications. However, due to the incomplete context, these approaches may face inevitable information loss in real workloads with more complex attention distributions.

Activation sharing: Another direction is sharing the intermediate activations to improve the attention calculation efficiency. Attention sharing approaches (Xiao et al., 2019; Wu et al., 2022; Li et al., 2021) observe the similarity among different layers’ attention matrix distribution and reuse these attention matrices to reduce the computation costs. Multi-query attention (MQA) (Shazeer, 2019) makes different heads share a single set of keys and values to reduce the memory bandwidth requirements in the incremental inference. Group-query attention (GQA) (Ainslie et al., 2023) relaxes the single set of keys and values restriction to multiple sets and each set is coupled with a group of queries. They have been successfully adopted by several recent public LLMs and shown their superior performance, including MQA-based models such as Falcon (ZXhang et al., 2023), PaLM (Chowdhery et al., 2022), ChatGLM2-6B (cha, 2023) and GQA-based models like LLaMA-2 (Touvron et al., 2023) and Mistral-7B (Jiang et al., 2023a).

Conditional computing: The sparsely-activated Mixture of Experts (MoE) (Shazeer et al., 2017; Csordás et al., 2023) paradigm partitions a model’s capacity across various “experts”, which are smaller neural networks, each specializing in different subsets of the data. It allows the system to only invoke the necessary experts for a given input based on certain routing mechanisms (Fedus et al., 2022; Lepikhin et al., 2020; Nie et al., 2021; Roller et al., 2021; Zhou et al., 2022a; Santos et al., 2023), rather than computing over the entire massive model, yielding computational and memory efficiency (Du et al., 2022). For example, TaskMoE (Kudugunta et al., 2021) illustrates that task-level routing enables model increase capacity compared with token-level counterpart, while improving the inference throughput. As LLMs continue to grow, the MoE architecture stands out as a promising avenue to ensure both scalability and efficiency for future LLMs. In the meanwhile, the dynamic nature of MoEs also demands special system optimization from both distributed communication (He et al., 2022; Nie et al., 2023; Rajbhandari et al., 2022; Hwang et al., 2023; Li et al., 2023b; Huang et al., 2023) and GPU kernel implementation (Gale et al., 2023; Zheng et al., 2023a) to facilitate MoE inference efficiency.

Recurrent unit: Although recurrent neural networks (RNN) (e.g., LSTM (Sak et al., 2014)) tend to struggle with capturing long-term dependencies in sequences (Khandelwal et al., 2018), there are still several approaches using recurrent units to replace Transformer modules and achieve linear computational and memory complexity during inference, such as RWKV (Peng et al., 2023a) and RetNet (Sun et al., 2023a). Specifically, unlike prior approaches, these recent explorations are mostly built on the linear attention (i.e., Linear Transformer (Katharopoulos et al., 2020), Attention Free Transformer (Zhai et al., 2021)) representation. After the reformation, they overcome the O(L2)\mathcal{O}(L^{2}) bottleneck of attention by modeling interactions between tokens with linear recurrence units (e.g., state space models (Gu et al., 2021; Mehta et al., 2022; Fu et al., 2022; Gu and Dao, 2023), LRU (Orvieto et al., 2023)), which are easier to maintain parallelizable training property. Their design is also composed of various position encoding modules (Su et al., 2021), exponential decay mechanisms (Oliva et al., 2017) and a stack of token-wise non-linear MLPs (Yu et al., 2022b; Tolstikhin et al., 2021) or GLUs (Dauphin et al., 2017) to improve the model representation capability. Recently, they have shown promising results on both model performance and computation efficiency. However, whether recurrent units can successfully replace Transformers for LLMs still remains an open problem (i.e., especially for long sequences).

1.3. Model Compression

Here, we delve into techniques for model compression, which aim to reduce the memory footprint and computational requirements of LLMs by creating more efficient and compact models without significant loss in performance.

Knowledge Distillation: One line of work is knowledge distillation, which trains a small student model with the supervision from a large teacher model. Most previous approaches in this direction are exploring white-box distillation (Sanh et al., 2019; Sun et al., 2019; Jiao et al., 2020; Wang et al., 2020b; Gu et al., 2023), which require accessing the entire teacher model parameters. Due to the arising of API-based LLM services (e.g., ChatGPT), several black-box distilled models attract lots of attention, such as Alpaca (Taori et al., 2023), Vicuna (Chiang et al., 2023), WizardLM (Xu et al., 2023c) and so on (Peng et al., 2023c; Zhu et al., 2023a). These models usually have fewer model parameters but have shown promising performance on various downstream tasks compared with the original LLMs (e.g., GPT-4 (OpenAI, 2023)).

Network pruning: Network pruning methods (Sanh et al., 2020; Michel et al., 2019; Sanh et al., 2020) have been extensively studied in the past few years but not all of them can be directly applied to LLMs. It is imperative to take into account the potentially exorbitant computational costs associated with retraining, as well as assess whether the pruning yields discernible gains in inference efficiency based on the underlying system’s implementation. Some recent approaches (Ma et al., 2023; Santacroce et al., 2023; Fan et al., 2019; Kurtic et al., 2023) apply structural pruning methods on LLMs, which removes entire structured LLM components, facilitating efficient GPU speedups. For example, Deja Vu (Liu et al., 2023f) cuts off specific attention heads and MLP parameters guided by the contextual sparsity hypothesis without modifying pre-trained models. There are also some recent advancements in unstructured methods (Frantar and Alistarh, 2023; Xu et al., 2023a; Sun et al., 2023b; Valicenti et al., 2023; Belcak and Wattenhofer, 2023), which usually achieve 50-60% sparsity for LLM compression. It is noteworthy that they can further generalize to semi-structured N:M sparsity (i.e., 2:4 and 4:8) (Mishra et al., 2021), leading to significant inference speedup with NVIDIA sparse tensor cores’ acceleration. LoSparse (Li et al., 2023d) and DSFormer (Chand et al., 2023) approximate model weights with a small dense and a sparse semi-structured matrix using low-rank factorization. Flash-LLM (Xia et al., 2023) relaxes this requirement by providing a memory-efficient SpMM implementation for unstructured pruning using tensor cores. PowerInfer (Song et al., 2023) assumes skew access of these sparsely-activated neurons and proposes a GPU-CPU hybrid inference engine, making GPU and CPU handle different neurons.

2. System Optimization

This section investigates LLM inference system optimization techniques to accelerate LLM inference without modifying the LLM computation semantics. The goal of this line of work is to improve the system efficiency by refining the underlying systems and frameworks used for large language model inference.

This section explores state-of-the-art low-bit quantization techniques that enable efficient representation of model weights and activations. By using fewer bits (i.e., less than 32) to represent numerical values, these methods significantly reduce memory consumption and accelerate inference on hardware platforms. One line of approach is to quantize LLM, and these quantization methods can be briefly categorized into two directions: Quantization-Aware Training (QAT) and Post-Training Quantization (PTQ) (Yao et al., 2023). PTQ reduces the computational precision of model weights (Frantar et al., 2022a, b; Dettmers et al., 2022; Lin et al., 2023; Dettmers et al., 2023b; Isik et al., 2023) and even activations (Yao et al., 2022; Xiao et al., 2022; Yuan et al., 2023) into either INT8 or INT4 by using custom CUDA kernels (Park et al., 2022; Li et al., 2023c) or compilations (Zhao et al., 2023) for efficiency benefits, such as W8A16 (i.e., INT8 weight-only quantization and FP16 or BF16 activations), W4A16 in GPTQ (Frantar et al., 2022a), W8A8 in SmoothQuant (Xiao et al., 2022) and W4A4 (Wu et al., 2023b). The evolution of hardware also meets these requirements. One supporting evidence is that NVIDIA’s recent architectures like Turing and Ampere have included INT8 and INT4 tensor cores, and the latest Hopper architecture has disabled INT4 support but introduced FP8 tensor cores for better numerical precision (e.g., H100 GPU can reach 60×\times TFLOPS for FP8 as opposed to FP32). Existing approaches usually adopt various quantization functions, including uniform methods (i.e., Round-to-Nearest) and non-uniform methods (Kim et al., 2023a). To relieve the performance loss from low-precision, QAT integrates quantization during model training (Liu et al., 2023e; Dettmers et al., 2023a). It is worth noting that due to challenges in the underlying system implementation, low-precision quantization methods may potentially result in slower inference speeds compared to conventional precision levels such as FP16 (Dettmers et al., 2022). While low-precision methods significantly reduce the resource requirements for model deployment, there is also research indicating that quantization methods can have a notable impact on the model’s inference performance due to the presence of scaling laws (Dettmers and Zettlemoyer, 2022). In addition, quantization has also been applied to context compression (e.g., CacheGen (Liu et al., 2023c)) and memory-efficient fine-tuning (e.g., QLoRA (Dettmers et al., 2023a), PEQA (Kim et al., 2023c)), resulting in lower memory consumption for LLM inference.

2.2. Parallel Computation

This section examines parallel computation strategies tailored for large language models. Leveraging parallel processing capabilities of modern hardware architectures, these methods distribute computation across multiple cores or devices, leading to substantial speedup during inference.

Model parallelism: Most model parallelism approaches are first proposed for distributed training of large-scale DNNs, especially for Transformer-based models. For example, tensor model parallelism (Shoeybi et al., 2019) (TP) splits the model layers (e.g., attention, FFN) into multiple pieces from internal dimensions (e.g., head, hidden) and deploys each on a separate device (e.g., GPU). It can significantly reduce inference latency through parallel computing, which is widely used for multiple GPUs within the same machine, especially for scenarios with high-speed NVLink connections. PaLM inference (Pope et al., 2023) extends TP on large-scale Transformer inference by involving 2D tensor parallelism (Van De Geijn and Watts, 1997) and claims lower theoretical communication complexity for large clusters (more than 256 devices). For multi-query attention with only one head for keys and values, it further involves data parallelism to the hybrid tensor partition strategy. Pipeline model parallelism (Narayanan et al., 2021) (PP) arranges the model layers in a sequence across multiple devices. Each device is responsible for a pipeline stage that consists of multiple consecutive model layers. While PP can significantly increase the number of inputs processed per unit of time (throughput), it doesn’t inherently decrease the time taken to process a single input from beginning to the end (latency) like TP. Sequence parallelism (SP) has various differentiated designs and implementations, but its key idea for LLM inference is to distribute the computational and storage load by splitting the processing of long sequences across multiple GPUs along the sequence length dimension (Liu et al., 2023g). Different parallelism techniques introduce varying degrees of communication overhead and computational latency (Isaev et al., 2023). To achieve optimal performance and resource utilization, automatic parallelism has been widely studied by prior approaches for distributed training (e.g., Alpa (Zheng et al., 2022), FlexFlow (Jia et al., 2019b; Unger et al., 2022), Galvatron (Miao et al., 2023b)). By replacing their cost model to fit the predictable runtime of auto-regressive inference of Transformer models like (Narayanan et al., 2023), it’s easy to apply previous automatic searching algorithms (e.g., dynamic programming, integer linear programming) to LLM serving (e.g., AlpaServe (Li et al., 2023e), FlexFlow-Serve (fle, 2023a), SpotServe (Miao et al., 2024)) and determine the most efficient parallelism strategy without manual intervention. There are also some approaches (Aminabadi et al., 2022; Sheng et al., 2023b; Miao et al., 2023a; Alizadeh et al., 2023; Guo et al., 2023) enabling offloading techniques to use larger but slower memory (e.g., CPU DRAM) to save model parameters and KV cache in addition to the limited device memory (e.g., GPU DRAM).

Decentralized inference: This line of approach involves a combination of model and data parallelism where multiple decentralized voluntary nodes collaborate to process data and infer outputs. This approach can be particularly useful in scenarios where hardware resources are geographically distributed. Inspired by crowdsourced computing, Petals (Borzunov et al., 2022) serves a BLOOM-176B model using collaborated commodity GPUs over the Internet. Decentralized inference opens up a new direction on unlocking the overlooked consumer-level GPUs for running LLMs, but also suffers from several practical challenges, such as device heterogeneity (Jiang et al., 2023d), limited computational and memory capacity, low-bandwidth network (Borzunov et al., 2023), fault tolerance and privacy protection (Tang et al., 2023).

2.3. Memory Management

Efficient memory management remains at the forefront of challenges in LLM serving, especially given the inherent memory-intensive nature of transformer architectures. With the growing need for long-sequence inference, the memory footprint of the KV cache stands out as a prime optimization target compared with model weights and the necessary workspace for other activations. As the KV cache memory grows and shrinks dynamically and unpredictably during incremental decoding, the naive approach (e.g., FasterTransformer) pre-allocates a contiguous piece of memory with a maximum sequence length assumption. It wastes memory severely for 1) input batches with varied request lengths and 2) complex decoding scenarios generating multiple output sequences in parallel (e.g., beam search, parallel decoding). vLLM (Kwon et al., 2023) proposes paged attention that partitions the KV cache into non-contiguous memory blocks and significantly improves the batch size as well as throughput. SpecInfer (Miao et al., 2023a) proposes tree attention and depth-first tree traversal to eliminate redundant KV cache allocation for multiple output sequences sharing the same prefix. LightLLM (lig, 2023) takes a more granular token-level memory management mechanism to further diminish memory usage. However, the overheads of such fragmented memory managing mechanisms pose new challenges. Especially for cases where other optimizations are employed to boost the batch size, these fine-grained memory management methods might offer only marginal throughput benefits while substantially amplifying the inference latency. It’s evident that memory reduction in LLM inference is intricately tied with other algorithmic innovations and system-level optimizations. While some might work well for specific workloads, they might counteract one another, leading to a degraded overall performance. Striking the right balance between memory efficiency and computational performance of LLM inference systems remains an open and pressing challenge in the field.

2.4. Request Scheduling

Efficiently scheduling incoming inference requests is crucial for optimizing LLM serving. This section reviews request scheduling algorithms that maximize resource utilization, guarantee response time within latency service level objective (SLO), and handle varying request loads effectively. Request scheduling for LLM serving shares commonalities with general ML serving techniques, as both aim to efficiently manage incoming requests and optimize resource utilization. These common aspects include dynamic batching (Ali et al., 2020), preemption (Han et al., 2022), priority (Ng et al., 2023), swapping (Bai et al., 2020), model selection (Gunasekaran et al., 2022), cost efficiency (Zhang et al., 2019), load balancing and resource allocation (Weng et al., 2022). However, LLM serving also introduces unique challenges due to its distinctive characteristics, such as the massive model size, iterative autoregressive decoding mechanism, unknown variable output length and state management for context information.

Early LLM serving systems (e.g., FasterTransformer over NVIDIA Triton) only support request-level scheduling which is similar to prior approaches. Orca (Yu et al., 2022a) first notices the gap between generative LLMs and the request-level scheduling of previous ML inference systems. Considering the variable output sequence length, it schedules the execution of the engine at the granularity of iteration with a first-come-first-serve (FCFS) order and enables batching a selected set of operations for better hardware utilization. Plenty of following approaches inherit the selective-batching and iteration-level scheduling policy, such as continuous batching in vLLM and RayLLM (ray, 2023) and in-flight batching in TensorRT-LLM (ten, 2023). Moreover, SpecInfer extends to speculative decoding by iteratively selecting a batch of requests to perform one iteration of speculative inference and verification. FastServe (Wu et al., 2023c) concentrates on the job completion time (JCT) and involves iteration-level preemption to prioritize requests with shorter input length, instead of FCFS. SARATHI (Agrawal et al., 2023) targets the pipeline bubbles in distributed inference caused by the initial iteration of varying length input requests. To saturate the GPU compute, it splits the input prompts into uniform chunks and piggybacks the chunk slot with other requests’ decoding iterations if possible, which is also adopted by DeepSpeed-FastGen called Dynamic SplitFuse (dee, 2023a). S3 (Jin et al., 2023) involves an output sequence length predictor and helps to schedule more concurrent requests within the GPU memory constraint for larger batch size and higher inference throughput.

2.5. Kernel Optimization

In this subsection, we delve into kernel-level optimizations, which target the performance of specific operations within the language model inference pipeline. These optimizations leverage hardware-specific features and software techniques to accelerate critical computation kernels.

Kernel fusion: To reduce overheads from kernel launching and memory accessing, kernel fusion is widely adapted by previous DNN frameworks and compilers. Since the backward computation is not required for LLM inference, more kernel fusion chances exist. Several contemporary Transformer inference engines (e.g., FasterTransformer (fas, 2021), TenTrans (Wu et al., 2021), TurboTransformers (Fang et al., 2021), LightSeq (Wang et al., 2020c), ByteTransformer (Zhai et al., 2023)) and compilers (e.g. Welder (Shi et al., 2023)) propose to fuse 1) GEMMs with the same shape (e.g., the three linear transformations for query, key and value) and 2) Add Bias with the other non-GEMM kernels, such as residual connection, layer normalization and activation functions (e.g., ReLU). Among these, the optimization of fused multi-head attention kernel has been extensively explored and will be discussed in the following aspect.

Tailored attention: To make the attention operations run efficiently on a GPU, customizing or tailoring the GPU kernels specifically for the attention calculation is crucial. For example, cuDNN has provided a fused multi-head attention kernel API (cud, 2023). Meanwhile, several implementations have been open-sourced for more performance gains. These can be roughly classified into two categories due to the special autoregressive decoding mechanism. One is for the first iteration (i.e., the initial/prefill/context/prompt phase), which processes all tokens from the input prompt in parallel. For example, xFormers (Lefaudeux et al., 2022) extends the online softmax trick (Rabe and Staats, 2021; Milakov and Gimelshein, 2018; Choi et al., 2022) to the whole attention calculation using CUTLASS (cut, 2023). The other is for the following iterations (i.e., the incremental/decode/generation phase) and the kernel only generates one output token per iteration. For autoregressive decoding, a common practice is to save the previously computed keys and values so that only a single query is required to compute when generating a new token instead of rerunning the entire sequence. The main direction of optimizations in this field is maximizing thread occupancy and minimizing the on-device high-bandwidth memory (HBM) access (i.e., using shared memory or registers (Chen et al., 2021a)). They usually parallelize across the batch size and number of heads dimension (e.g., FasterTransformer) to distribute workloads. Some further enable parallelizing the sequence length dimension by partitioning the KV cache into chunks but require reducing the chunk-wise results at last, such as FlashDecoding (Tri Dao, [n. d.]). A subsequent work FlashDecoding++ (Hong et al., 2023) removes such synchronization for partial softmax by introducing a unified maximum value known in advance. It is necessary to select the appropriate parallel dimension based on the workloads for better thread utilization.

Sampling optimization: The sampling algorithm selection can greatly influence the LLM generation quality. The default greedy sampling always picks the token with the highest probability. Parallel sampling techniques, such as beam search, decode the approximate optimal sequences efficiently by maintaining a fixed number (i.e., beam width) of top-scoring sequences every iteration. A variety of stochastic sampling techniques (e.g., top-kk (Fan et al., 2018), top-pp (Holtzman et al., 2019), temperature controlling (Keskar et al., 2019)) have been propose to introduce randomness for more diverse outputs. However, they are still suffering from several practical system challenges. One is the increased memory pressure from redundant KV cache (§3.2.3), and another is the sampling efficiency issue attributed by the large vocabulary of LLM (i.e., tens of thousands). For example, LightSeq (Wang et al., 2020c) provides an efficient hierarchical implementation that divides the vocabulary into kk groups, retrieves candidates within each group using a few GPU instructions and then re-ranks these candidates to obtain the top-kk tokens.

Variable sequence length: Another unique challenge of LLM inference is that the sequences can vary in both input length and output length, and the latter is unknown in advance. One way to speed up inference is to process multiple sequences in a batch at once (§3.2.4). However, when a batch of sequences has variable input lengths, padding is often used to make them all the same length for batch processing, wasting computational and memory resources. To alleviate some of these inefficiencies, various strategies can be employed. Packing technique (pac, 2020; Zhai et al., 2023) stores the sequences into a continuous memory space without padding and only unpacks before attention calculation. Ragged tensor (Fegade et al., 2022) further supports computation with minimal padding using compiler-generated kernels. Bucketing the sequence into a smaller computation granularity (e.g., chunks (Du et al., 2023)) is also a possible solution to alleviate memory usage of padding tokens. Due to the mixed execution of the initial phase and incremental phase, bucketing input prompts (Agrawal et al., 2023) also brings new challenges to the memory management and request scheduling (§ 3.2.4).

Automatic compilation: Most existing LLM inference systems utilize vendor-specific libraries as their backend, such as cuBLAS, cuDNN and CUTLASS, which provide optimized kernel implementations. To further improve the inference efficiency, they also take great efforts on optimizing manually-written kernels for specific LLM operators (e.g., attention) over NVIDIA GPUs. Despite of these work, the trend of using automated DNN compilers still exists, such as TVM (i.e., Unity (Sampson et al., 2022), Relax (Lai et al., 2023) and TensorIR (Feng et al., 2023; Ye et al., 2023)), MLIR (Katel et al., 2022), JAX (Frostig et al., 2018), OpenAI Triton (Tillet et al., 2019), TASO (Jia et al., 2019a) and TorchInductor (Wu, 2023). The compilation approach can help discover potentially more efficient operator implementations (e.g., expression derivation (Zheng et al., 2023b)), and more importantly, facilitate adaptation to alternative hardware platforms, including mobile and edge devices, CPUs, DL accelerators, and other types of GPUs (e.g., AMD GPUs and Apple M2 Ultra).

Software Frameworks

Generative LLM serving requires a full stack of optimizations and many recent works have started to develop software frameworks to provide efficient LLM inference deployment service. In the following, we revisit these systems and investigate a comprehensive analysis of several representative open-sourced GPU-based LLM serving systems in Table 2. The analysis does not contain some popular related projects, including 1) specialized solutions for other hardware (e.g., PopTransformer (pop, 2023), CTranslate2 (ctr, 2023), lammap.cpp and ggml (ggm, 2023)) and 2) deployment solutions built on top of the other systems, like OpenLLM (ope, 2023) (vLLM), xinference (xin, 2023) (ggml + vLLM + xFormers), LMDeploy (lmd, 2023) (FasterTransformer), gpt-fast (pyt, 2023) (PyTorch), DeepSpeed-MII and DeepSpeed-FastGen (dee, 2023b) (DeepSpeed-Inference), and RayLLM and RayServe (ray, 2023) (vLLM).

We compare these state-of-the-art LLM serving systems and summarize their differences in several aspects. First, most of these systems support tensor parallelism to enable multi-GPU inference and improve the system performance. And some of them future support pipeline parallelism or offloading to support inference over multi-node or resource-constrained environments individually. Second, partial systems learn from Orca and implement the iteration-level scheduling. Third, we investigate the attention kernels of these systems and introduce their implementations in terms of the initial and incremental phases respectively. For the initial phase, they usually adapt a batched general matrix multiply (GEMM) approach (e.g., cuBLAS, torch, Relay) and some utilize the online softmax trick to reduce HBM access (e.g., Flash-attention, xFormers). The incremental phase is more challenging because the per-token generation scheme results in lower computational intensity. To improve the GPU utilization, FasterTransformer manually fuses the attention calculations (e.g., linear projection, positional bias, dot product, softmax, etc) into a single high-performance kernel template and involves several kernel optimization techniques, such as caching with shard memory, warp-shuffle instruction for reduction, half matrix multiplication and accumulation (HMMA) with tensor core and multiple-precision support. FlexFlow-Serve enables speculative decoding and provides a tree-based parallel decoding kernel to verify the speculated tokens from multiple sequences (i.e., from multiple small models or different beams or parallel sampling) with zero-memory redundancy and maximum thread parallelism. vLLM extends the fused mutli-head attention (MHA) kernel from from FasterTransformer by partitioning the KV cache into pages to eliminate redundant memory usage, especially for parallel sampling scenarios. LightLLM takes a follow-up approach by partitioning the KV cache into more fine-grained token-wise pieces.

Note that, there still remain some other notable aspects that are not covered by the above discussions. For example, even for the most popular Flash and Paged attention kernels, they are usually implemented in different ways across these systems. TGI directly imports the original Flash/Paged attention libraries, LightLLM adopts kernels implemented by OpenAI Triton, MLC-LLM generates kernels by TVM, and TensorRT-LLM modifies from FasterTransformer’s fused attention kernel to support paged attention. Another example is about the input-aware kernel selection. For the initial phase, TensorRT-LLM selects from cuBLAS and Flash attention based on the context length. Besides the attention calculation, for the linear projection operators, there is also a recent trend of replacing GEMM with general matrix-vector product (GEMV) to handle the cases of small batch size (i.e., 1) more efficiently. And these systems also have many other different features, such as programming language (i.e., C++, Python), low-precision support (i.e., FP16, INT8), supported hardware and models. In summary, these different choices of design and implementation are largely determined by their prioritized optimization target. For example, vLLM proposes paged attention to improve the batch size for higher throughput (TptT_{pt}), while FlexFlow-Serve leverages SpecInfer to accelerate decoding for lower latency (LatL_{at}). Basically, low latency and high throughput are dual optimization targets in LLM serving systems, representing complementary but often conflicting objectives, necessitating a balanced strategy to optimize the trade-off between rapid response for individual tasks and maximizing the volume of tasks processed over a specified time frame. Some recent studies (Databricks, 2023) further decompose the response latency by TTFT+TPOT ×\times output sequence length, where TTFT represents Time To First Token and TPOT represents Time Per Output Token. The former is driven by the initial phase processing speed while the latter directly depends on per-iteration execution time during incremental decoding. Distinguishing these two metrics is beneficial to LLM service providers, leading to different system design choices and user experience (e.g., faster application responsiveness (Liu et al., 2023c), longer prompts (dee, 2023a)). Besides, reducing the monetary cost is also an important and practical objective for the design and implementation of some LLM serving systems (Miao et al., 2024). Although it unlikely to have a one-size-fits-all solution, we believe that future LLM serving systems will continually integrate these differentiated features, thereby continuously improving system efficiency and hardware utilization.

Benchmarks

Building a comprehensive and reproducible benchmark for comparing the performance of various LLM serving system like MLPerf (Reddi et al., 2020) is a critical endeavor for both academic and industrial communities in this field. It will not only help LLM users select the right system solutions but also encourage researchers and developers to keep pace with the advanced optimizations. Unfortunately, despite of some prior reports (ham, 2023; llm, 2023), up to this point, the community has not yet launched a convincing enough benchmark that takes into account all influencing factors. This is mainly because of the numerous evaluation settings, including model configuration, hardware environment, and request load, among others. Testing under a limited number of setting combinations cannot yield conclusions with credibility. For example, certain system optimization techniques can only achieve performance advantages under high or low load conditions, and conversely, they might even be detrimental. Besides, when measuring inference latency, how to exclude additional overheads not related to GPU inference (such as request scheduling overhead, inherent network latency, etc.) due to differences in system design is also a challenging topic. Additionally, a fair benchmark test needs to consider the strict alignment of model output content, which is often overlooked in many tests.

Connection with other surveys

Our survey on efficient generative LLM serving and inference complements and extends the scope of existing literature in the field, while maintaining a distinct focus. Among the related works, (Kim et al., 2023b) comes closest in subject matter exploring the design of more general Transformer models and domain-specific accelerators. However, our survey differentiates itself by focusing specifically on generative LLM serving, a nuanced area that has not been the central focus of other studies. Moreover, some studies delve into experimental investigations of LLM inference efficiency on GPUs (Narayanan et al., 2023; Zhang et al., 2023b) and novel accelerators (Emani et al., 2023), offering valuable empirical insights that are directly relevant to our focus on serving efficiency. Additionally, LLMCarbon (Faiz et al., 2023) addresses an increasingly important aspect of LLM deployment – its environmental impact (e.g., carbon footprints). While our survey’s primary focus is efficiency from a performance standpoint, the environmental lens provided by such studies is undeniably relevant and respected in our broader discussion. Some surveys and benchmarks (Jaiswal et al., 2023) offer valuable insights into model compression (Zhu et al., 2023b; Gupta and Agrawal, 2022; Zhu et al., 2023b; Treviso et al., 2023) and quantization (Yao et al., 2023; Gholami et al., 2022). These studies lay a groundwork that indirectly supports our exploration of related directions. Some studies (Muhlgay et al., 2023; Dalvi et al., 2023) provide essential context for understanding LLM effectiveness (e.g., accuracy, perplexity, factuality and so on), which is beyond the scope of this survey. Our survey also acknowledges the contributions of prior surveys (Ben-Nun and Hoefler, 2019; Mayer and Jacobsen, 2020) focusing on distributed training of large-scale DNN models, as they inform the backdrop against which LLM serving must be considered. In essence, our survey situates itself amidst a diverse array of studies, drawing from and contributing to a more holistic understanding of LLM serving efficiency, including both algorithmic innovations and system optimizations. By integrating insights from these various areas, we aim to provide a nuanced and comprehensive overview of the latest advancements and challenges in the field.

Future Direction

As we stand at the forefront of LLM advancements, it becomes increasingly important to not only understand the current state of these technologies but also to anticipate and shape their future trajectory. Particularly in the realm of generative LLM serving, there is a vast landscape of unexplored possibilities and emerging challenges. The rapid evolution of this field necessitates a forward-looking approach, where identifying potential avenues for innovation and improvement is crucial. This foresight not only prepares us to adapt to upcoming technological shifts but also guides the research community toward addressing the most pertinent and impactful areas. In this context, we outline several promising directions for future research and development, each offering the potential to significantly enhance the efficiency of serving generative LLMs.

Future progress in enhancing generative LLM serving efficiency could be significantly driven by the development and refinement of specialized hardware accelerators, complemented by a co-design approach that aligns hardware and software optimizations. For instance, integrating memory closer to processing units or optimizing chip architectures to better align with the data flow of LLM algorithms can lead to substantial reductions in latency and energy consumption. This approach has been exemplified in recent GPU advancements, like NVIDIA’s Hopper architecture (nvh, 2022), which demonstrates improvements in HBM and SRAM capacity, memory bandwidth, computing units and bisection bandwidth, directly benefiting the processing of LLMs. Continued innovation in this area could involve designing hardware that is inherently tuned to the computational patterns of generative LLMs, such as optimizing for the specific demands of attention mechanisms and tensor operations that are prevalent in these models, eventually influencing the design and implementation of LLM serving systems.

The development of more efficient decoding algorithms could substantially improve serving efficiency. Motivated by the demand for more resource-efficient ways to utilize the vast knowledge encapsulated within LLMs, future work could explore alternative approaches to the traditional auto-regressive methods and unlock the generation speed for real-time applications while maintaining the decoding quality. One promising direction is generalized speculative inference as it enables preserving the same generation quality. Specifically, the small speculative model can be generalized to any other forms of methods that can generate draft tokens more efficiently than LLMs, such as knowledge retriever and user-defined functions (Miao et al., 2023a; Yang et al., 2023a). For example, some subsequent works arose recently, replacing the draft model with early exiting (Yang et al., 2023b; Zhang et al., 2023e; Bae et al., 2023; Hooper et al., 2023) or non-autoregressive decoding (Ge et al., 2022; Fu et al., 2023). In summary, the development of efficient decoding algorithms like speculative decoding coupled with the underlying system optimizations represents a significant opportunity to enhance the serving efficiency of generative LLMs.

As the application of LLMs continues to expand into more sophisticated scenarios, the demand for processing longer contexts or sequences is steadily growing. Serving LLMs with long-sequence workloads requires resolving the challenges from both the algorithm and system sides. In terms of LLMs, they often suffer from length generalization failure when sequences get longer than what was observed during training (Press et al., 2021) even enabling relative positional encoding (Chen et al., 2023b) or after fine-tuning on longer corpora (Bai et al., 2023). Even for some models that claim to support ultra-long contexts, studies have found that they encounter a situation of “loss in the middle” (Liu et al., 2023d). Current approaches attempt to alleviate such limitations by reducing the computational sequence length while preserving relevant information, such as retrieval augmentation (Xu et al., 2023b), sequence compression (Jiang et al., 2023c) and caching (Gim et al., 2023). For the LLM serving systems, longer sequence brings critical challenges, including more memory consumption and access of KV cache and quadratic increasing computational complexity of self-attention.

Although Transformer models and self-attention mechanisms currently dominate the landscape of LLMs, exploring alternative architectures is a promising direction for future research. The field of DL has historically seen a constant alternation of dominant architectures, with each new paradigm shift bringing about significant advancements. Given this trend, it’s important to consider other architectural approaches that could offer distinct advantages, especially for improved computational efficiency. For instance, some recent studies explore attention-free methods (Bozic et al., 2023), using pure MLP (Multi-Layer Perceptron) architectures to replace attention mechanisms. The evolution of DNN model architecture is not only a natural progression, but also a necessary exploration to uncover more efficient and effective ways of structuring LLMs.

As the application of LLMs expands, a crucial future direction involves exploring and optimizing their deployment across various complex environments. This exploration goes beyond traditional cloud-based deployments to include scenarios like edge computing, hybrid computing (combining cloud and edge computing), decentralized computing, and the utilization of more affordable resources like spot instances. Each of these environments presents unique challenges and opportunities for LLM serving. For instance, edge computing allows for faster response times and reduced bandwidth usage by processing data closer to the source, but it also poses challenges in terms of limited computational resources and storage capacity. Hybrid computing (Qualcomm, 2023) offers a balanced approach but requires advanced management to distribute computational tasks efficiently. Decentralized computing presents a promising avenue for crowdsourcing computational resources, but it also brings additional considerations regarding data privacy and security (Zhang et al., 2023a; Lu et al., 2023). LLM serving over preemptive resources (Miao et al., 2024) can significantly reduce monetary costs but requires fault tolerance mechanisms to handle their inherent unpredictability and variability, ensuring consistent performance and system reliability. Successfully navigating the challenges from these complex environments will be key for more robust, scalable, and efficient LLM applications.

The diverse application-specific requirements create a wide range of innovative LLM serving optimization opportunities, such as parameter-efficient fine-tuning (Zhou et al., 2022b; Sheng et al., 2023a; Chen et al., 2023c), retrieval from external vector storage (Borgeaud et al., 2022), online learning and knowledge updates, multi-modal workloads, and chaining together different LLMs’ capabilities (Wu et al., 2023a). These unique challenges also demand automatic and smooth integration of LLM serving techniques into existing IT infrastructures by extending the optimization space to the whole LLM lifetime, including data acquisition and processing, AutoML (Tornede et al., 2023) and model management (Nagrecha and Kumar, 2023), resource allocations, and performance monitoring.

Conclusion

Efficient LLM serving is a fundamental step towards democratizing access to advanced AI technologies. This survey aims to provide researchers, practitioners, and developers with a comprehensive understanding of the existing methodologies, enabling them to make informed decisions when deploying LLMs in real-world environments. By consolidating the latest research findings on algorithms and systems, this survey paper hopes to accelerate progress and foster innovation in the pursuit of highly efficient LLM serving solutions.

References