Intriguing Properties of Quantization at Scale

Arash Ahmadian, Saurabh Dash, Hongyu Chen, Bharat Venkitesh, Stephen Gou, Phil Blunsom, Ahmet Üstün, Sara Hooker

Introduction

The push for ever larger language models (LLMs) has been driven by a strong correlation between performance and the number of parameters (Chowdhery et al., 2022; Zhang et al., 2022; Kaplan et al., 2020). This has led to new breakthroughs in downstream performance, but has also posed new challenges in making these models accessible. Larger models incur higher memory and latency because of the requirement to store many more model weights and the optimizer in fixed memory (Dehghani et al., 2021; Treviso et al., 2022). Due to the massive size of state-of-art LLMs, inference often requires hosting across multiple machines which limits the practical usability of such models.

To address this, much research has focused on compression techniques such as quantization, which reduces the number of bits needed to represent each learned parameter in a model (Gholami et al., 2021). Quantization techniques are widely used in smaller model regimes – quantizing weights stored as 32-bit or 16-bit floating-point numbers to 8-bit integers (INT8) produces large reductions in memory and latency.

However, at scale simple quantization techniques have been shown to lead to a pronounced degradation in performance (Xiao et al., 2022). This trade-off has been attributed by several recent works (Dettmers et al., 2022; Zeng et al., 2022; Xiao et al., 2022; Bondarenko et al., 2021) to emergent outlier dimensions—scaling transformer-based architecture results in large activation magnitude outliers which are concentrated along a few hidden dimensions. To remedy this degradation, mixed-precision solutions have been proposed that handle the outlier activations separately (Dettmers et al., 2022). While effective at preventing degradation, these specialized techniques pose significant latency overhead which negates some of the benefits to memory (Wu et al., 2020).

The difficulties of quantizing at scale prompt us to ask: are emergent properties due to nature or nurture? Recent work introduces intriguing and somewhat contradictory answers to this question: Models like OPT-175B (Zhang et al., 2022) and FairSeq (Artetxe et al., 2022) exhibit pronounced sensitivity to post-training quantization and require complex mixed-precision decomposition quantization methods (Dettmers et al., 2022; Wei et al., 2022b; Bondarenko et al., 2021; Luo et al., 2020; Zeng et al., 2022). On the contrary, BLOOM-176B (Scao et al., 2022) is easier to quantize with a simple quantization recipe and a relatively small performance drop (Frantar et al., 2022; Xiao et al., 2022). Zeng et al. (2022) hypothesize that the observed difference in weight distribution characteristics may be due to the difference in optimization choices made during pre-training.

In this work, we seek to reconcile these observations. We posit that it is possible to optimize for a quantization friendly training recipe that suppresses large activation magnitude outliers. This leads to a distribution of activations and weights that are more amenable to simple INT8 quantization recipes and does not necessitate the need for complex and inefficient mixed-precision computations. Our results show that we can introduce simple INT8 post-training quantization with negligible impact on performance due to choices we make during the pre-training stage. As shown in Figure 1, across 8 zero-shot downstream tasks, our models do not present any significant performance drop, having only 0.24% average degradation in a 52 billion parameter model.

In summary, our contributions are as follows:

We conduct a controlled large scale study – we maintain the same architecture and vary key optimization choices such as weight decay, gradient clipping, dropout and precision of training representation. We present results across models varying from 410 million to 52 billion parameters, with each experiment variant trained from random initialization. While this requires a compute intensive set-up, it allows us to rigorously disentangle what factors actually influence sensitivity to quantization.

We show that reoccurring activation outliers are not a universal emergent property of LLMs at scale and can be avoided at scales as large as 52B given the right optimization choices. Our 52B parameter model shows only 0.26% performance degradation across 8 tasks with INT8 PTQ quantization of both activations and weights.

We contribute a fine-grained analysis of activations and weights, and show that several key weight and activation characteristics may explain the difference in sensitivity between our robust models and models like OPT which have been shown to have pronounced sensitivity to quantitization at scale. We hope these insights help guide future model design and pre-training strategies.

Background

Quantization refers to compressing weights and activations of a neural network into lower-bit representations. Here, our focus is on one-shot post-training quantization (PTQ) (Xiao et al., 2022; Dettmers et al., 2022), which quantizes the network post-training without additional finetuning steps for calibration. Given the complexities of successfully training a large language model (Zhang et al., 2022; Rae et al., 2021), PTQ methods are extremely attractive as these techniques require the least modification to pre-trained parameters, compared to quantization-aware training methods (Zafrir et al., 2019; Krishnamoorthi, 2018) or quantization-aware finetuning (Park et al., 2022b; Yao et al., 2022; Frantar et al., 2022; Zhuo et al., 2022; Li et al., 2021; Hubara et al., 2020; Nagel et al., 2020) which both require updates to model weights. We include more detail about each of these broad groups of techniques in Appendix C.

To-date PTQ techniques that quantize both the activations and weights of a language model have proven extremely challenging at large scale (>>6B parameters) leading to pronounced drops in performance. Instead, less aggressive techniques have been used such as weight-only quantization (Gerganov, 2023; Frantar et al., 2022; Zeng et al., 2022) that leaves the activations in higher precision or quantization with mixed-precision decomposition (Dettmers et al., 2022) which decomposes the matrix multiplication to compute a small fraction of elements at a higher precision (FP16) while the bulk of the computations is performed at low precision (INT8).

Weight-only quantization brings speedup to inference by reducing the amount of data movement. However, as large language models are scaled, progressively they become compute-bound and the improvements due to weight-only quantization stagnate. While mixed-precision decomposition approaches have theoretical latency benefits due to the bulk of the computation being performed at lower precision, in practice without specialized hardware (Dash et al., 2022; Dash & Mukhopadhyay, 2020; Hooker, 2021), GPU kernels, or additional kernel calls to prepare the inputs and weights for mixed-precision computation, the projected benefits cannot be realized (Dettmers et al., 2022). To further realize latency gains, we need to quantize both weights and activations into 8-bit integers (INT8) to utilize specialized integer INT8 GEMM kernels, which are supported by a wide range of hardware (e.g., NVIDIA GPUs, Intel CPUs, Qualcomm DSPs, etc.) In addition, weight and activations quantization further enables compressing key-value cache to INT8. Since key-value cache takes up a significant part of the GPU memory during inference (Sheng et al., 2023), weight and activations quantization further contributes to memory saving and high throughput inference.

Where ⊙\odot denotes broadcastable matrix multiplication. The quantization and dequantization steps for the above do not add much memory overhead compared to quantizing with a single scaling constant for each weight or activation tensor, while significantly increasing the representation power of INT8.

Methodology and Experimental Setup

Our goal is to understand whether sensitivity to widely used quantization techniques is inherently an emergent property at scale or due to optimization choices made during pre-training. Recent work has presented seemingly contradictory empirical findings – some models such as OPT-175B show pronounced sensitivity at scale to post-training quantization while other models such as BLOOM-176B are relatively robust to post-training quantization.

These models differ in numerous ways such as architectural differences, pre-training data, training infrastructure, and finally optimization choices,making it challenging to attribute differences in quantization performance. To rigorously isolate what choices result in sensitivity to quantization, we measure the impact of optimization choices within a tightly controlled experimental setup – training the same large scale model architecture from random initialization while rigorously varying only key aspects of the optimization procedure. Each optimization choice is evaluated in two ways: we measure the resulting degradation after PTQ in zero-shot downstream performance and then analyze the model weights and feature activations to understand how the characteristics at scale impact quantization performance.

Training multiple multi-billion parameter size language models is extremely expensive – a single 52B language model takes roughly 20 days of training with 2048 TPU cores.22We include more details about the hardware and training requirements in Section 3.2. Therefore, we first conduct our controlled experiments on 410M and 6B models using early checkpoints and then validate the results at scale, by fully training 6B, 13B, and 52B parameter size models with our most quantization friendly training recipe. In practice, we found performance at early checkpoints predictive of fully trained model performance.

We briefly describe each of the axes of variations below:

Weight decay Weight decay is widely used to impede over-fitting by penalizing large magnitude weights (Goodfellow et al., 2016). We experiment with a range of weight decay values {0.001, 0.01, 0.1}.

Gradient clipping Gradient clipping rescales the norm of the gradient vector if it exceeds the threshold (Pascanu et al., 2013). It is widely used in LLMs to prevent exploding gradients and accelerate training convergence (Du et al., 2021; Zhang et al., 2022). We experiment with a gradient norm threshold of 1 as well as training without gradient clipping.

Dropout Dropout is a widely used regularization technique that drops neurons with a probability of pp during training (Srivastava et al., 2014)(Hinton et al., 2012). We apply dropout to the output of the self-attention block and the feed-forward block before the corresponding residual connection as described in Vaswani et al. (2017), but we do not use a dropout for the input embeddings. We experiment with {0, 0.1, 0.4, 0.8} dropout probabilities.

Half-precision data type: bf16 vs fp16 Training neural networks in mixed-precision is a common technique to reduce the memory requirement and improves training time while often achieving comparable performance to full-precision training (Micikevicius et al., 2017). In this technique, a copy of weights is stored in full-precision (fp32) whereas the forward and backward passes are done using half-precision in either float16 (fp16) or bfloat16 (bf16) (Kalamkar et al., 2019; Dean et al., 2012; Abadi et al., 2015).

We experiment with fp16 and bf16. Furthermore, for each half-precision data type, we vary weight decay values of (0.1, 0.01) to observe whether the effect of the half-precision data type is exasperated with a smaller weight decay value of 0.01.

2 Experimental Setup

Model We train autoregressive decoder-only Transformer models (Liu et al., 2018) with a standard language modeling objective. Given an input sequence of S=S= [s1,⋯ ,st]\left[s_{1},\cdots,s_{t}\right], a language model with parameters θ\theta trained to minimizes the following negative log likelihood:

Our language models follow the traditional GPT style architecture reported in Radford et al. . Different from the Radford et al. , we do not share the same weight matrix for input and output embedding layers. Instead, we learn separate input and output projections to enable higher model parallelism. We train models with parameter sizes ranging from 410 million to 52 billion. All models have a maximum sequence length of 2048 tokens. We use SentencePiece (Kudo & Richardson, 2018) tokenizer with a vocabulary of 51200 to tokenize the text.

Training details We pre-train models using a mixture of datasets from Common Crawl and C4 (Raffel et al., 2020) with AdamW (Loshchilov & Hutter, 2019) optimizer and a batch size of 256. We use a cosine learning rate scheduler with 1500 warm-up steps. We use GeLU activations (Hendrycks & Gimpel, 2016). All the models are trained with mixed-precision, i.e. forward and backward passes are computed in bf16 or fp16 half-precision format, but the model parameters are stored in fp32 in the distributed optimizer state (Rajbhandari et al., 2020). For half-precision fp16, we experimented with a variant by switching layernorm arithmetic from fp16 to fp32 for better numeric stability (Micikevicius et al., 2017). Without layernorm arithmetic being in fp32, the training was unable to converge.

To avoid an exorbitant computational cost given each variant requires training from scratch and we are evaluating very large scale models, we first iterated on 410M models but observed very minimal degradation. As a result, we scaled to 6B parameters which is the scale at which we present our results. Following the analysis, we validate our findings by scaling the most PTQ friendly optimization choices to 13B and 52B models until convergence. The optimal training hyper-parameters for PTQ are provided in Appendix A.1.

Infrastructure We use TPU-v4 chips (Jouppi et al., 2017) to train, and Nvidia A100 GPUs to evaluate our models. All models are trained using the FAX (Yoo et al., 2022) framework which enables efficient model and data parallelism. It takes approximately 72 hours on 128 cores to train a 6B parameter model for 75000 steps.

Evaluation We evaluate each model variant on Copa (test and dev set) (Wang et al., 2019), HellaSwag (Zellers et al., 2019), PIQAValidation (Bisk et al., 2020), StoryCloze (Mostafazadeh et al., 2016), WinoGrande (Sakaguchi et al., 2019), Paralex (Fader et al., 2013), and LAMBADA (Paperno et al., 2016). All evaluations were done in a zero-shot setting. In total, we benchmark on 88 tasks comprised of Multiple choice (MC) completion, MC Co-referencing, Generation, and Question Answering (QA) types. Details of our evaluation suite are given in Appendix A.2. For each experimental variant, we report the average percent performance difference and the pre-quantization and post-quantization performances across tasks. The percent degradation is calculated by normalizing each task’s absolute degradation by the corresponding pre-quantization performance.

Results and Discussion

For each experimental axis, we train the corresponding variants to a maximum of 75000 steps. Below we present a breakdown of the degradation results and analysis for each experimental axis. All variants with the exception of dropout=0.8, had similar pre-quantization performance. This is important, as we are interested in comparing optimization choices that still result in models of comparable quality, but differing sensitivities to post-training quantization. Refer to Appendix A.3 for the per-task breakdown of results.

Weight Decay As can be seen in Figure 2(a), we observe that a higher level of weight decay during pre-training improves post-training quantization performance. We do not use gradient clipping in these experiments to isolate the impact of weight decay. A larger weight decay value (0.1 vs 0.001) results in better post-training performance (0.09% vs 1.36% degradation). Furthermore, as shown in Figure 3, combining lower weight decay with fp16 can further amplify sensitivity to post-training quantization. A small weight decay value (0.01) can cause higher performance degradation in post-quantization after training (1.73%).

Dropout and Gradient Clipping In Figure 2(b) we observe that higher levels of dropout correspond to sharper degradation in post-training quantization. Note that Pdropout=0.8P_{dropout}=0.8 unsurprisingly leads to a poor absolute downstream performance, however, it helps establish a clear trend given other data points. Figure 2(c) shows the relative quantization degradation for models with and without gradient clipping. When varying gradient clipping, a control weight decay value of 0.001 is used to minimize the impact of weight decay. As seen in the figure, gradient clipping shows a positive impact on the quantization performance, improving robustness to post-training quantization. This suggests that gradient clipping to an extent counteracts the effects of a small weight-decay value which would otherwise lead to higher quantization degradation.

Half-precision: bf16 vs fp16 Figure 3 shows the quantization degradation and absolute performance for fp16 and bf16 for 6B parameter models. Training with fp16 leads to higher quantization degradation than bf16. We relate the degradation results to the numerical stability in training. fp16 format uses a smaller range for the exponent than bf16. While bf16 uses 8 bits for the exponent, fp16 only uses 5. Most floating point formats also have denormalized numbers which allow for a soft underflow. This can get exponentially closer to 0.0f for each additional bit in the mantissa. This makes underflow more of a concern for floating point formats.

Notably, we observe that our findings provide insights into the quantization robustness of BLOOM-176B, which to our knowledge is the only open-source LLM (with more than 50B parameters) and a decoder-only block architecture, that is trained with bf16; compared to simliar sized fp16 trained models such as OPT-175B (Xiao et al., 2022).

Using Early Checkpoints to Infer Converged Model Behavior Given the considerable computational cost of our experiments, it is valuable to explore whether converged model behavior can be inferred from checkpoints from early snapshots of training. In Figure 3, we plot the relative post-training quantization degradation given checkpoints trained with different levels of precision at different steps during pre-training. We observe that quantization degradation increases with step count, but the relative impact of varying the bit representation emerges early in training which confirms the main trend. Interestingly, fp16 (wd=0.01) variant exhibits high quantization degradation in the starting phase training as early as 15000 steps.

Scaling Insights to 52B scale To validate our experimental findings at scale and with fully trained models, we pre-train 410M, 6B, 13B, and 52B parameter models using the best optimization choices with respect to robustness in post-training quantization: weight decay of 0.1, no dropout, gradient clipping of 1, and bf16 as the half-precision format. Figure 1 shows mean zero-shot accuracy for the non-quantized and quantized model using INT8 weights and activations. Compared with OPT models (Zhang et al., 2022), our fully-trained models are significantly more robust to post-training quantization starting from 6B parameter size. Our largest scale model with 52B parameters, shows a 0.08% improvement in average performance across the evaluation suite and only 0.01% degradation across LAMBADA, HellaSwag, and PIQA where OPT-66B which is the closest OPT model in terms of size, has an extreme drop of ∼\sim42% as reported in Dettmers et al. (2022).

To evaluate our models’ performance under a different quantization recipe in addition to INT8, we also test 4-bit integer (INT4) column-wise weight-only quantization, similar to Du et al. (2021). Our 52B parameter model exhibited only a 3.6% relative drop in mean zero-shot performance across the 8 evaluation tasks. It is worth noting that this quantization scheme does not require any fine-tuning or optimization and hence these results highlight the high robustness of our trained models. In comparison, when applying the same quantization scheme to BLOOM, BLOOM-176B and BLOOM-7B show 29.5% and 18.7% degradation respectively on LAMBADA (Du et al., 2021), while our 52B model only has 8.6% degradation.

Weight and Activation Analysis

Our results in Section 4 find that sensitivity to quantization at scale is not an inherent emergent property. Rather, it can be avoided at scales as large as 52B given the right optimization choices. In this section, we perform a fine-grained analysis of activations and weights to understand how the trained distribution of our models differs from models like OPT that are far more sensitive to quantization. For all the metrics proposed below, we include the complete analysis for all layers in Appendix B.

Activations As a first step, we analyze input activations of the attention projection (attn-kqv-proj) as it is the earliest point for INT8 multiplication in a decoder block. Here, we measure root-mean-square error RMSE(X,X^)\text{RMSE}(\mathbf{X},\mathbf{\hat{X}}) where X^\mathbf{\hat{X}} denotes the de-quantized activations. Additionally, we report the mean standard deviation of the input activations measured per token. While RMSE(X,X^)\text{RMSE}(\mathbf{X},\mathbf{\hat{X}}) directly indicates the quantization error, standard deviation (STD) has been shown to be closely related to the expected quantization error of a normally distributed input (Kuzmin et al., 2022). Figure 5 compares the bf16 and fp16 variants. We observe that the RMSE and STD of fp16 are far higher than the bf16 variant – the RMSE for the fp16 variant is 6.9x the RMSE for the bf16 variant. This difference is even more pronounced if we compare our model to the OPT: the RMSE and STD of the OPT are 27.7x and 1.8x higher respectively relative to our model (Figure 5; Bottom row).

In the bottom row of Figure 5, we also compare STD(g\mathbf{g}) of our model relative to OPT and BLOOM. We observe that even the BLOOM model that is relatively robust to quantization has a far larger STD(g\mathbf{g}) than our model with a multiplier of 5x. Interestingly, we find that OPT-6B layernorm gain parameters are all set to 1.0 while biases varied as expected. Hence, given the gain parameters appear to be hardcoded, the STD(g\mathbf{g}) of the OPT model is 0. We were not able to find any mention of such design decision either in Zhang et al. (2022) or the github repository: https://github.com/facebookresearch/metaseq.

Attention Projection Weights Finally, we compare the weight distribution of attn-kqv-proj layers. As seen in Figure 4, the fp16 variant has a significantly wider distribution compared to bf16. Additionally, inspired by Lin et al. (2019), we use spectral norm to measure the maximum degree of noise amplification for each token activation.

where σmax\sigma_{max} is the largest singular value of W\mathbf{W}.

As seen in Figure 5, we observe that the spectral norm of the fp16 variant is 4x higher than the bf16. In addition, on the bottom row of Figure 5, we observe that both BLOOM and our model have generally lower spectral norm than OPT 6B that is far more sensitive to quantization.

Discussion In addition to the metrics proposed above, we also sought to incorporate recent work that proposes to use the number of outlier dimensions as a proxy measure to understand sensitivity to quantization degradation (Dettmers et al., 2022). Activation feature dimensions are classified as outlier dimensions when the activation values (denoted as α\alpha) are greater than 6.0 at more than 25% of layers and 6% of tokens. However, we find that a threshold of 6.0 is too high to classify a feature dimension as an outlier for all the variants we consider. After correspondence with the authors, we also explored various adaptations of this outlier detection recipe presented in Dettmers et al. (2022) to make it generalizable. However, we did not observe a clear correlation between these measures and sensitivity to quantization. We refer to Appendix B.3 for detailed treatment of these replication efforts. Due to these observations, we hope that the metrics we proposed and evaluated in this Section will help further discussion about useful proxy metrics for guiding pre-training optimization choices to improve robustness to quantization.

Related Work

Challenges of Quantization at Scale Recently, there have been several studies to characterize the emergence of outliers at scale, and relate this to the difficulties in post-training quantization of both weights and activations (Dettmers et al., 2022; Wei et al., 2022b; Puccetti et al., 2022). Dettmers et al. (2022) depict a phenomenon of emerging outliers by observing that large outlier dimensions systematically emerge at a certain scale (6.7B parameters) which hamper quantization attempts. Extreme outliers at scale was also empirically confirmed in follow-up works (Zeng et al., 2022; Xiao et al., 2022). The causes of outliers have also been the subject of recent work. Puccetti et al. (2022) observe that in Masked Language Models (MLMs) the magnitude of hidden state coefficients corresponding to outlier dimensions correlates with the frequency of encoded tokens in pre-training data. Wei et al. (2022b) observe that LayerNorm scaling (g\mathbf{g}) amplifies the outliers and can be suppressed using a modified LayerNorm and token-wise clipping. Wortsman et al. (2023) consider large-scale vision-language models and show that quantization techniques are more stable if the network is trained and initialized so that large feature magnitudes are discouraged.

Most mitigation strategies to quantize in the presence of outliers has required more complex quantization techniques. For example, Dettmers et al. (2022) propose selective mixed-precision computation by only computing the outliers at higher precision. However, such a setup proves difficult to map to hardware, limiting the inference speedup. Xiao et al. (2022) propose to smoothen out these outliers by migrating some of the activation variances into the model weights with appropriate scaling. Although the authors demonstrate the ability of this framework to scale to large models, additional rescaling is required for activations which leads to additional latency overhead without specialized kernels. Another limitation of Xiao et al. (2022) is that it relies on the assumption that outliers exist in activations, and that weights can bear additional outliers and still be easy to quantize.

Our work is the first to show that outliers are not inherent to scaling large language models. Rather than an emerging property, they are a result of particular training methods. Compared to previous methods using extensive quantization schemes with custom kernels (Dettmers et al., 2022), our work applies PTQ using simple, one-shot linear weight and activation quantizations which can take advantage of NVIDIA-provided CUTLASS kernels, leading to a significant decrease in latency and memory footprint.

Conclusion

We present a rigorous study of the effect of how various optimization choices affect INT8 PTQ with the goal of reconciling the recent contradictory observations regarding emergent properties in Large Language Models. We show that regularization directly impacts PTQ performance and that higher levels of regularization through common techniques such as weight-decay, and gradient-clipping leads to lower post-training quantization degradation. We further demonstrate that the choice of half-precision training data type has a significant impact on PTQ performance – emergent features are significantly less pronounced when training with bf16.

Broader Impact Our work serves as a useful counter-example to scholarship which has advanced the notion that certain properties depend only on model scale (Wei et al., 2022a). Rather, our results support the conclusion that optimization choices play a large role in whether emergent properties are present. We believe there is more work to be done here. We also hope that the insights gained from our work illustrate the significant impact the underlying hardware can have on PTQ. Currently, bf16 training is possible on TPUs and only very recently introduced to A100 & H100 GPUs. Finally, we belive our results present an impactful formula for training models which are inherently easier to quantize at scale, making these models more accessible for deploying in a variety of deployment environments.

Limitations We do not vary the architectural design and training objective in our experiments given our goal of a controlled experimental set-up and the large computational cost of each variant. We leave exploring the impact of different training objectives and architecture design choices to future work.

Acknowledgement

We thank João Araújo, Milad Alizadeh and other colleagues in Cohere & Cohere For AI for helpful feedback and support. We also thank Tim Dettmers for assisting in replicating the outlier dimension definition and results in int8.LLM().

References

Appendix

Appendix A Extended Results & Architecture Details

A.2 Evaluation Suite

Below is a detailed breakdown of the evaluation suite we evaluate our models with.

A.3 Task Result Breakdown

Data type PIQA HellaSwag WinoGrande LAMBADA Copa Copa100 StoryCloze Paralex Average Average % Diff wd=0.1 (gc=none) FP16 75.35 60.58 55.41 56.59 73.20 71.00 77.02 58.75 65.99 -0.09 INT8 75.35 60.36 55.25 57.69 73.20 70.00 76.70 58.62 65.90 wd=0.01 (gc=none) FP16 75.03 60.55 55.25 59.01 72.00 69.00 77.53 58.99 65.92 -0.26 INT8 74.48 60.10 55.49 60.14 72.40 67.00 77.21 58.87 65.71 wd=0.001 (gc=none) FP16 75.90 60.71 55.80 58.16 73.00 71.00 76.64 58.50 66.21 -1.36 INT8 75.63 60.52 54.78 55.15 71.80 70.00 77.21 57.96 65.38 dtype=bf16 (wd=0.1) FP16 75.68 60.92 55.96 57.60 71.40 71.00 75.81 59.05 65.93 0.32 INT8 75.95 60.70 56.83 58.32 71.40 71.00 75.62 59.06 66.11 dtype=fp16 (wd=0.1) FP16 73.83 59.96 55.96 56.14 72.00 67.00 75.94 58.33 64.89 -0.76 INT8 74.16 59.75 55.64 56.05 71.60 64.00 75.88 58.12 64.40 dtype=bf16 (wd=0.01) FP16 74.97 60.79 55.49 57.52 72.00 68.00 75.88 58.74 65.42 0.77 INT8 75.08 60.51 55.88 59.31 72.60 69.00 76.00 58.84 65.90 dtype=fp16 (wd=0.01) FP16 74.81 58.11 54.93 57.31 70.20 71.00 74.67 58.25 64.91 -1.73 INT8 73.61 56.90 54.22 53.02 71.60 69.00 74.67 57.92 63.87 gc=1.0 (wd=0.001) FP16 74.65 60.03 54.78 59.01 71.60 67.00 76.96 58.92 65.37 0.41 INT8 74.92 59.91 54.62 59.69 72.20 68.00 76.96 58.90 65.65 gc=none (wd=0.001) FP16 75.90 60.71 55.80 58.16 73.00 71.00 76.64 58.50 66.21 -1.36 INT8 75.63 60.52 54.78 55.15 71.80 70.00 77.21 57.96 65.38 dropout=0.0 FP16 75.68 60.92 55.96 57.60 71.40 71.00 75.81 59.05 65.93 0.32 INT8 75.95 60.70 56.83 58.32 71.40 71.00 75.62 59.06 66.11 dropout=0.1 FP16 74.76 58.87 54.38 57.23 71.60 68.00 76.45 58.36 64.96 0.31 INT8 74.27 58.70 54.85 58.35 71.80 68.00 76.96 58.18 65.14 dropout=0.4 FP16 74.92 55.80 54.70 58.98 71.00 66.00 74.03 57.91 64.17 -0.27 INT8 74.76 55.77 54.14 59.69 69.40 66.00 74.09 57.95 63.98 dropout=0.8 FP16 67.79 30.87 50.51 37.12 67.00 65.00 61.94 29.16 51.17 -0.57 INT8 67.79 30.75 50.12 36.52 66.60 65.00 61.74 28.91 50.93

Appendix B Extended Weight & Activation Analysis

B.2 Layers Analysis

B.3 Outlier analysis

We classify a feature dimension as an outlier dimension if the same dimension is classified as an outlier across 20 random samples from the C4 validation set (Raffel et al., 2020). These samples are fixed when experimenting with different outlier definitions. As mentioned in Section 4, we experimented with the definition outlined in Dettmers et al. (2022) where a hidden feature dimension is classified as outlier dimension if the activation magnitudes (denoted as α\alpha) are greater than 6.0 (α>6\alpha>6) at more than 25% of layers and 6% of tokens. However, we find 0 outlier dimensions across all of our 6B variants and fully trained models using this definition. As a result, following the methodology outlined in Dettmers et al. (2022), we manually searched for the lowest threshold such that only one dimension is classified as outlier in our smallest (410M) fully trained model. As shown in Table 5, even classifying outlier dimensions using the searched threshold of 4.2 resulted in most variants having 0 outlier dimensions. We further experimented with not fixing the threshold and using z-score outlier detection i.e. classify a feature as a high-magnitude if α>Cσ+μ\alpha>C\sigma+\mu, where σ\sigma and μ\mu denote sample standard deviation and mean. We were not able to establish clear trends with this method either as shown in Table 5.

Appendix C Extended Literature Review

The need for compression techniques that scale to large language model settings has become increasingly urgent with larger and larger models (Treviso et al., 2022; Yao et al., 2023). There has been a renewed focus on efficiency techniques (Gale et al., 2019; Ogueji et al., 2022; Ahia et al., 2021). Quantization as a form of model compression of large language models has become increasingly relevant as a way to minimize memory requirements and minimize compute intensity (Dettmers et al., 2022; Xiao et al., 2022; Frantar et al., 2022; Park et al., 2022a; Kim et al., ).

Model Efficiency at Inference Time Research in model compression mostly falls in the categories of quantization techniques (Jacob et al., 2018; Courbariaux et al., 2014; Hubara et al., 2016; Gupta et al., 2015), efforts to start with a network that is more compact with fewer parameters, layers or computations (architecture design) (Howard et al., 2017; Iandola et al., 2016; Kumar et al., 2017), student networks with fewer parameters that learn from a larger teacher model (model distillation) (Hinton et al., 2015) and finally pruning by setting a subset of weights or filters to zero (Louizos et al., 2017; Wen et al., 2016; LeCun et al., 1990; Hassibi et al., 1993a; Ström, 1997; Hassibi et al., 1993b; See et al., 2016; Narang et al., 2017; Frantar & Alistarh, 2023; Sanh et al., 2020). Often, a combination of compression methods might be applied. For example, pruning might be combined with other efficiency-improving methods, e.g. quantization or faster search algorithms.

Quantization Techniques Quantization can be used to speed up inference and relax hardware requirements, as has been shown for e.g., 8-bit (Quinn & Ballesteros, 2018), 4-bit (Aji & Heafield, 2020) and recently also below 3-bit quantization (Park et al., 2022a) of neural machine translation models. Park et al. (2022a) utilize non-uniform quantization to achieve high compression ratios. However, this method requires the use of specialized kernels for compressed (2-bit/4-bit) weights and floating point activations and involve finding the binary representations using expensive iterative search or QAT. Similar to Park et al. (2022a), Frantar et al. (2022) also demonstrate compression of model parameters to 3 or 4-bit precision allowing inference off a single A100 GPU. However, they also perform weight-only quantization limiting the speedup as the activations are kept at higher precision (FP16).

Quantization reduces the number of bits needed to represent model weights which minimizes both the memory and latency required to serve a model.

Often, the goal is to quantize the bit representation while preserving equivalent performance. Quantization approaches can be broadly categorized into:

Quantization-aware training (QAT) (Zafrir et al., 2019; Krishnamoorthi, 2018) – Quantization-aware training (QAT) involves pre-training with simulated quantization, enabling parameters to adjust to lower precision grids. This requires estimating the derivative of non-differentiable quantization operators, performing full backpropagation throughout the entire model, and training with the entire training dataset. However, this method can be computationally expensive, particularly for large language models.

Quantization-aware finetuning (Yao et al., 2022; Frantar et al., 2022; Zhuo et al., 2022; Li et al., 2021; Hubara et al., 2020; Nagel et al., 2020) (QAF) is a more efficient approach that utilizes a pretrained model and a small subset of training data (i.e., hundreds of samples) to optimize performance under quantization. By simulating quantization and optimizing a small range of parameters at a time, no backpropagation is needed while the quantization loss can be reduced.

One-shot post-training quantization (PTQ) (Xiao et al., 2022; Dettmers et al., 2022) unlike QAT and QAF, does not involve optimization. Instead, it directly maps data from a high precision range to a low precision range based on a hand-picked mapping function.

Given the complexities of successfully training a large language model (Zhang et al., 2022; Rae et al., 2021), post-training quantization (PTQ) methods are extremely attractive as these techniques require the least modification to pretrained parameters. This is the focus of our exploration in this work.

Below section introduces widely used quantization methods and provides context about the differences between these methods. The quantization strategy for weights and activations can be broadly classified into three categories:

Weight-only quantization has proven extremely effective in making large language models accessible by enabling inference in a resource-constrained environment while maintaining the FP16 model quality (Gerganov, 2023; Frantar et al., 2022; Zeng et al., 2022). Weight-only quantization provides improvements in latency due to a reduction in time taken for parameter fetching from GPU global memory, however, the actual Matrix-Matrix multiplication (GEMM) operations are carried out at higher precision in FP16 - allowing modest gains on platforms without dedicated lower-precision GEMM operations support.

C.1.2 Weight and Activation Quantization

As large language models are scaled, progressively they become compute-bound and the improvements due to weight-only quantization stagnate. However, in this regime, using efficient kernels that leverage specialized lower-precision cores in modern GPUs to directly perform the actual Matrix-Matrix multiplication operation at lower precision enables large latency gains - due to the increased throughput of INT8 tensor cores over FP16 Tensor Cores (Nvidia, ). As this quantization technique scales the best, this is going to be our focus in this work.

To-date quantization of both the activations and weights of very large models (>>6.7B parameters) has proven challenging - leading to a large drop in performance.

C.1.3 Quantization by Mixed-Precision Decomposition

In the quantization strategies mentioned above, even though the various weights and activations might be stored in different precisions; all the computations in a single operation are carried out at the same precision (FP16 or INT8). In contrast, LLM.int8() (Dettmers et al., 2022) proposes to decompose the matrix multiplication to compute a small fraction of elements at a higher precision (FP16) while the bulk of the computations is performed at low precision (INT8). This approach has a similar footprint to that of weight-only quantization but practical latency gains are limited or potentially worse. While this approach has theoretical latency benefits due to the bulk of the computation being performed at lower precision, in practice without specialized hardware (Dash et al., 2022; Dash & Mukhopadhyay, 2020), the lack of specialized kernels on GPUs and additional kernel calls required to ready the inputs and weights for mixed-precision computation negates the projected benefits. In this work we focus on exploring optimization choices which mitigate quantization trade-offs for both weight-only quantization and the far more challenging weight and activation quantization.