I-BERT: Integer-only BERT Quantization
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney, Kurt Keutzer
Introduction
The recent Transformer based Neural Network (NN) models , pre-trained from large unlabeled data (e.g., BERT , RoBERTa , and the GPT family ), have achieved a significant accuracy improvement when fine-tuned on a wide range of Natural Language Processing (NLP) tasks such as sentence classification and question answering . Despite the state-of-the-art results in various NLP tasks, pre-trained Transformer models are generally orders of magnitude larger than prior models. For example, the BERT-Large model contains 340M parameters. Much larger Transformer models have been introduced in the past few years, with even more parameters . Efficient deployment of these models has become a major challenge, even in data centers, due to limited resources (energy, memory footprint, and compute) and the need for real-time inference. Obviously, these challenges are greater for edge devices, where the compute and energy resources are more constrained.
One promising method to tackle this challenge is quantization , a procedure which compresses NN models into smaller size by representing parameters and/or activations with low bit precision, e.g., 8-bit integer (INT8) instead of 32-bit floating point (FP32). Quantization reduces memory footprint by storing parameters/activations in low precision. With the recent integer-only quantization methods, one can also benefit from faster inference speed by using low precision integer multiplication and accumulation, instead of floating point arithmetic. However, previous quantization schemes for Transformer based models use simulated quantization (aka fake quantization), where all or part of operations in the inference (e.g., GELU , Softmax, and Layer Normalization ) are carried out with floating point arithmetic . This approach has multiple drawbacks for deployment in real edge application scenarios. Most importantly, the resulting NN models cannot be deployed on neural accelerators or popular edge processors that do not support floating point arithmetic. For instance, the recent server class of Turing Tensor Cores have added high throughput integer logic that are faster than single/half-precision. Similarly, some of the edge processor cores in ARM Cortex-M family for embedded systems only contain integer arithmetic units, and they can only support NN deployment with the integer-only kernels . Moreover, one has to consider that compared to the integer-only inference, the approaches that use floating point arithmetic are inferior in latency and power efficiency. For chip designers wishing to support BERT-like models, adding floating point arithmetic logic occupies larger die area on a chip, as compared to integer arithmetic logic. Thus, the complete removal of floating point arithmetic for inference could have a major impact on designing applications, software, and hardware for efficient inference at the edge .
While prior work has shown the feasibility of integer-only inference , these approaches have only focused on models in computer vision with simple CNN layers, Batch Normalization (BatchNorm) , and ReLU activations. These are all linear or piece-wise linear operators. Due to the non-linear operations used in Transformer architecture, e.g., GELU, Softmax, and Layer Normalization (LayerNorm), these methods cannot be applied to Transformer based models. Unlike ReLU, computing GELU and Softmax with integer-only arithmetic is not straightforward, due to their non-linearity. Furthermore, unlike BatchNorm whose parameters/statistics can be fused into the previous convolutional layer in inference, LayerNorm requires the dynamic computation of the square root of the variance for each input. This cannot be naïvely computed with integer-only arithmetic. Another challenge is that processing GELU, Softmax, and LayerNorm with low precision can result in signifciant accuracy degradation . For these reasons, other quantization methods such as keep these operations in FP32 precision.
In this work, we propose I-BERT to address these challenges. I-BERT incorporates a series of novel integer-only quantization scheme for Transformer based models. Specifically, our contributions are:
We propose new kernels for the efficient and accurate integer-only computation of GELU and Softmax. In particular, we approximate GELU and Softmax with light-weight second-order polynomials, which can be evaluated with integer-only arithmetic. We utilize different techniques to improve the approximation error, and achieve a maximum error of for GELU, and for Softmax. See § 3.4 and 3.5 for details.
For LayerNorm, we perform integer-only computation by leveraging a known algorithm for integer calculation of square root . See § 3.6 for details.
We use these approximations of GELU, Softmax, and LayerNorm to design integer-only quantization for Transformer based models. Specifically, we process Embedding and matrix multiplication (MatMul) with INT8 multiplication and INT32 accumulation. The following non-linear operations (GELU, Softmax, and LayerNorm) are then calculated on the INT32 accumulated result and then requantized back to INT8. We represent all parameters and activations in the entire computational graph with integers, and we never cast them into floating point. See Fig. 1 (right) for a schematic description.
We apply I-BERT to RoBERTa-Base/Large, and we evaluate their accuracy on the GLUE downstream tasks. I-BERT achieves similar results as compared to full-precision baseline. Specifically, I-BERT outperforms the baseline by 0.3 and 0.5 on the GLUE downstream tasks for RoBERTa-Base and RoBERTa-Large, respectively. See Tab. 2 in § 4.1 for details.
We deploy INT8 BERT models with the integer-only kernels for non-linear operations on a T4 GPU using TensorRT . We show that INT8 inference achieves up to 4 speedup as compared to FP32 inference. See Tab. 3 in § 4.2 for details.
Related Work
Efficient Neural Network. There are several different approaches to reduce the memory footprint, latency, and power of modern NN architectures. These techniques can be broadly categorized into: (1) pruning ; (2) knowledge distillation ; (3) efficient neural architecture design ; (4) hardware-aware NN co-design ; and (5) quantization.
Here, we only focus on quantization and briefly discuss the related work.
Quantization. For quantization, the parameters and/or activations are represented with low bit precision . While this line of research mostly focuses on CNN models, there have been recent attempts to introduce quantization techniques into Transformer based models as well. For example, and propose an 8-bit quantization scheme for Transformer based models and compress the model size up to 25% of the original size. Another work applies uniform and mixed-precision to quantize BERT model, where a second-order sensitivity method is used for the mixed-precision setting. quantizes a different subset of weights in each training iteration to make models more robust to quantization. Recently, there have been attempts to quantize BERT with even lower precision. presents a 3/4-bit centroid-based quantization method that does not require fine-tuning. leverage knowledge distillation to ternarize/binarize weights. combines knowledge distillation and learned step size quantization method to achieve up to 2-bit quantization of BERT.
However, to the best of our knowledge, all of the prior quantization work on Transformer based models use simulated quantization (aka fake quantization), where all or part of operations are performed with floating point arithmetic. This requires the quantized parameters and/or activations to be dequantized back to FP32 for the floating point operations. For example, perform the entire inference using floating point arithmetic, as schematically shown in Fig. 1 (left). While attempt to process Embedding and MatMul efficiently with integer arithmetic, they keep the remaining operations (i.e., GELU, Softmax, and LayerNorm) in FP32, as illustrated in Fig. 1 (middle). However, our method I-BERT uses integer-only quantization for the entire inference process—i.e., without any floating point arithmetic and without any dequantization during the entire inference. This is illustrated in Fig. 1 (right). This allows more efficient hardware deployment on specialized accelerators or integer-only processors as well as faster and less energy consuming inference. While we focus on uniform quantization, our method is complementary to other mixed and/or low-precision methods, and can be deployed for those settings as well.
To briefly discuss, there are also several quantization works for computer vision. introduces an integer-only quantization scheme for popular CNN models, by replacing all floating point operations (e.g., convolution, MatMul, and ReLU) with integer operations. Similarly, the recent work of extends this approach to low precision and mixed precision dyadic quantization, which is an extension of integer-only quantization where no integer division is used. However, both of these works are limited to CNN models that only contain linear and piece-wise linear operators, and they cannot be applied to Transformer based models with non-linear operators, e.g., GELU, Softmax, and LayerNorm. Our work aims to address this limitation by extending the integer-only scheme to the Transformer based models without accuracy drop.
Methodology
Under uniform symmetric quantization scheme, a real number is uniformly mapped to an integer value , where specifies the quantization bit precision. The formal definition is:
2 Non-linear Functions with Integer-only Arithmetic
To address this challenge, we approximate non-linear activation functions, GELU and Softmax, with polynomials that can be computed with integer-only arithmetic. Computing polynomials consists of only addition and multiplication, which can be performed with integer arithmetic. As such, if we can find good polynomial approximations to these operations, then we can perform the entire inference with integer-only arithmetic. For instance, a second-order polynomial represented as can be efficiently calculated with integer-only arithmetic as shown in Alg. 1.In Alg. 1, means the floor function. Note that, , , and can be pre-computed under static quantization. That is to say, there is no floating point calculation, e.g., of , in inference.
3 Polynomial Approximation of Non-linear Functions
There is a large body of work on approximating a function with a polynomial . We use a class of interpolating polynomials, where we are given the function value for a set of different data points , and we seek to find a polynomial of degree at most that exactly matches the function value at these points. It is known that there exists a unique polynomial of degree at most that passes through all the data points . We denote this polynomial by , defined as:
Interestingly for our problem, we have two knobs to change to find the best polynomial approximation. Since we know the actual target function and can query its exact value for any input, we can choose the interpolating point to be any point on the function. The second knob is to choose the degree of the polynomial. While choosing a high-order polynomial results in smaller error (see Appendix B), there are two problems with this. First, high-order polynomials have higher computational and memory overhead. Second, it is challenging to evaluate them with low-precision integer-only arithmetic, as overflow can happen when multiplying integer values. For every multiplication, we need to use double bit-precision to avoid overflow. As such, the challenge is to find a good low-order polynomial that can closely approximate the non-linear functions used in Transformers. This is what we discuss next, for GELU and Softmax, in § 3.4 and 3.5, respectively, where we show that one can get a close approximation by using only a second-order polynomial.
4 Integer-only GELU
GELU is a non-linear activation function used in Transformer models, defined as:
where is the Sigmoid function. This approximation, however, is not a viable solution for integer-only quantization, as the Sigmoid itself is another non-linear function which requires floating point arithmetic. One way to address this is to approximate Sigmoid with the so-called hard Sigmoid (h-Sigmoid) proposed by (designed in the context of efficient computer vision models) to obtain an integer-only approximation for GELU:
We refer to this approximation as h-GELU. Although h-GELU can be computed with integer arithmetic, we observed that replacing GELU with h-GELU in Transformers results in a significant accuracy drop. This is due to the large gap between h-GELU and GELU as depicted in Tab. 1.Later in our ablation study, we show this can lead to accuracy degradation of up to 2.2 percentages, as reported in Tab. 4. Figure 2 (left) also shows the noticeable gap between those two functions.
A simple way to address the above problem is to use polynomials to approximate GELU, by solving the following optimization problem:
Algorithm 2 summarizes the integer-only computation of GELU using i-GELU. We illustrate the behaviour of i-GELU in Fig. 2 (left). As one can see, i-GELU closely approximates GELU, particularly around the origin. We also report the approximation error of i-GELU along with h-GELU in Tab. 1, where i-GELU has an average error of and a maximum error of . This is more accurate than h-GELU whose average and maximum errors are and , respectively. Also, i-GELU even slightly outperforms the Sigmoid based approximation of Eq. 5, but without using any floating point arithmetic. Note that computing the Sigmoid requires floating point. Later in the results section, we show that this improved approximation, actually results in better accuracy of i-GELU as compared to h-GELU (see Tab. 4).
5 Integer-only Softmax
Softmax normalizes an input vector and maps it to a probability distribution:
Approximating the Softmax layer with integer arithmetic is quite challenging, as the exponential function used in Softmax is unbounded and changes rapidly. As such, prior Transformer quantization techniques treat this layer using floating point arithmetic. Some prior work have proposed look up tables with interpolation , but as before we avoid look up tables and strive for a pure arithmetic based approximation. In addition, although proposes polynomial approximation methods for the exponential function, it uses significantly high-degree polynomials, and is only applicable on a limited finite domain.
Similar to GELU, we cannot use a high-order polynomial, but even using such polynomial is ineffective to approximate the exponential function in Softmax. However, it is possible to address problem by limiting the approximation range of Softmax. First, we subtract the maximum value from the input to the exponential for numerical stability:
where >> is the bit shifting operation. As a result, we only need to approximate the exponential function in the compact interval of . This is a much smaller range as compared to the domain of all real numbers. Interestingly, a variant of this method was used in the Itanium 2 machine from HP , but with a look up table for evaluating .
We use a second-order polynomial to approximate the exponential function in this range. To find the coefficients of the polynomial, we minimize the L2 distance from exponential function in the interval of . This results in the following approximation:
Substituting the exponential term in Eq. 12 with this polynomial results in i-exp:
6 Integer-only LayerNorm
LayerNorm is commonly used in Transformers and involves several non-linear operations, such as division, square, and square root. This operation is used for normalizing the input activation across the channel dimension. The normalization process is described as:
Here, and are the mean and standard deviation of the input across the channel dimension. One subtle challenge here is that the input statistics (i.e., and ) change rapidly for NLP tasks, and these values need to be calculated dynamically during runtime. While computing is straightforward, evaluating requires the square-root function.
The square-root function can be efficiently evaluated with integer-only arithmetic through an iterative algorithm proposed in , as described in Alg. 4. Given any non-negative integer input , this algorithm iteratively searches for the exact value of based on Newton’s Method and only requires integer arithmetic. This algorithm is computationally lightweight, as it converges within at most four iterations for any INT32 inputs and each iteration consists only of one integer division, one integer addition, and one bit-shifting operation. The rest of the the non-linear operations in LayerNorm such as division and square are straightforwardly computed with integer arithmetic.
Results
In this section, we first measure the accuracy of I-BERT using the General Language Understanding Evaluation (GLUE) benchmark (§ 4.1). Then, we discuss the latency speedup of I-BERT using direct hardware deployment and compare it with pure FP32 model (§ 4.2). Finally, we conduct ablation studies to showcase the effectiveness of our integer-only approximation methods (§ 4.3).
We implement I-BERT on the RoBERTa model using . For the integer-only implementation, we replace all the floating point operations in the original model with the corresponding integer-only operations that were discussed in § 3. In particular, we perform MatMul and Embedding with INT8 precision, and the non-linear operations with INT32 precision, as using INT32 for computing these operations has little overhead. See § C.1 for implementation details. For each of the GLUE downstream tasks, we train both FP32 baseline and integer-only I-BERT models, and evaluate the accuracy on the development set. See Appendix C.2 and C.3 for training and evaluation details. While we only test RoBERTa-Base/Large, our method is not restricted to RoBERTa. The integer-only approximations can be performed for any NN models including Transformers that uses similar non-linear operations.
The integer-only quantization results for RoBERTa-Base/Large are presented in Tab. 2. As one can see, I-BERT consistently achieves comparable or slightly higher accuracy than baseline. For RoBERTa-Base, I-BERT achieves higher accuracy for all cases (up to 1.4 for RTE), except for MNLI-m, QQP, and STS-B tasks, where we observe a small accuracy degradation up to 0.3. We observe a similar behaviour on the RoBERTa-Large model, where I-BERT matches or outperforms the baseline accuracy for all the downstream tasks. On average, I-BERT outperforms the baseline by 0.3/0.5 for RoBERTa-Base/Large, respectively.
2 Latency Evaluation
We evaluate the latency speedup of INT8 inference of I-BERT, by direct deployment on a Tesla T4 GPU with Turing Tensor Cores that supports accelerated INT8 execution. Although T4 GPU is not a pure integer-only hardware, we select it as our target device due to its extensive software support , and in particular Nvidia’s TensorRT library . Furthermore, as we do not exploit any T4-specific exclusive features or requirements, our work can be extensively deployed on other hardware as well. See § C.4 for the detailed environment setup. For evaluation, we implement two variants of BERT-Base/Large: (1) pure FP32 models using naïve FP32 kernels for non-linear operations; and (2) quantized INT8 models using customized kernels for the non-linear operations. The customized kernels compute GELU, Softmax, and LayerNorm based on the integer-only methods described in § 3. We measure the inference latency for different sequence lengths (128 and 256) and batch sizes (1, 2, 4, and 8).
Table 3 shows the inference latency speedup of INT8 models with respect to FP32 models. As one can see, the INT8 inference of I-BERT is on average 3.08 and 3.56 faster than pure FP32 inference for BERT-Base and BERT-Large, respectively, achieving up to 4.00 speedup. The result implies that, when deployed on specialized hardware that supports efficient integer computations, I-BERT can achieve significant speedup as compared to FP32 models. Further speedups are possible with NVIDIA’s custom Transformer plugins which fuse the multi-head attention and Softmax layers (see § C.4).
While the greatest value of our work will become evident when our approach enables quantization on lower-end microprocessors without floating-point hardware, this demonstration must wait for improved software support for implementing quantized NN models on those processors. In the meantime, we believe the promise of our approach is illustrated by these latency reductions shown above.
3 Ablation Studies
Here, we perform an ablation study to show the benefit of i-GELU as compared to other approximation methods for GELU, and in particular h-GELU in Eq. 6. For comparison, we implement two variants of I-BERT by replacing i-GELU with GELU and h-GELU, respectively. The former is the exact computation of GELU with floating point arithmetic, and the later is another integer-only approximation method for GELU (see § 3). We use RoBERTa-Large model as baseline along with the QNLI, SST-2, MPRC, and RTE tasks. All models are trained and fine-tuned according to the procedure described in § 4.1, and the final accuracies are reported in Tab. 4.
As one can see, replacing GELU with h-GELU approximation results in accuracy degradation for all downstream tasks except for MRPC. Accuracy drops by 0.5 on average and up to 1.1 for RTE task. Although accuracy slightly improves for MRPC, the amount of increase is smaller than replacing GELU with i-GELU. This empirically demonstrates that h-GELU is not sufficiently tight enough to approximate GELU well. Approximating GELU with i-GELU results in strictly better accuracy for all four downstream tasks than h-GELU. In particular, i-GELU outperforms h-GELU by 0.7 on average, and it achieves comparable or slightly better result to the non-approximated full-precision GELU. i-GELU also performs better than GELU, which is quite interesting, but at this time, we do not have an explanation for this behaviour.
Conclusions
We have proposed I-BERT, a novel integer-only quantization scheme for Transformers, where the entire inference is performed with pure integer arithmetic. Key elements of I-BERT are approximation methods for nonlinear operations such as GELU, Softmax, and LayerNorm, which enable their approximation with integer computation. We empirically evaluated I-BERT on RoBERTa-Base/Large models, where our quantization method improves the average GLUE score by 0.3/0.5 points as comapred to baseline. Furthermore, we directly deployed the quantized models and measured the end-to-end inference latency, showing that I-BERT can achieve up to 4.00 speedup on a Tesla T4 GPU as compared to floating point baseline. As part of future work, one could consider using our approximation to improve the training speed as well. For instance, one could consider replacing GELU with i-GELU during training. Also, further studies are needed to evaluate the performance benefit of i-GELU as compared to GELU.
Acknowledgments
The UC Berkeley team acknowledges gracious support from Intel corporation, Intel VLAB team, Google Cloud, Google TRC team, and Nvidia, as well as valuable feedback from Prof. Dave Patterson, and Prof. Joseph Gonzalez. Amir Gholami was supported through a gracious fund from Samsung SAIT. Michael W. Mahoney would also like to acknowledge the UC Berkeley CLTC, ARO, NSF, and ONR. Our conclusions do not necessarily reflect the position or the policy of our sponsors, and no official endorsement should be inferred.
References
Appendix A Quantization Methods
A.2 Static and Dynamic Quantization
There is a subtle but important factor to consider when computing the scaling factor, . Computing this scaling factor requires determining the range of parameters/activations (i.e., parameter in Eq. 1). Since the model parameters are fixed during inference, their range and the corresponding scaling factor can be precomputed. However, activations vary across different inputs, and thus their range varies. One way to address this issue is to use dynamic quantization, where the activation range and the scaling factor are calculated during inference. However, computing the range of activation is costly as it requires a scan over the entire data and often results in significant overhead. Static quantization avoids this runtime computation by precomputing a fixed range based on the statistics of activations during training, and then uses that fixed range during inference. As such, it does not have the runtime overhead of computing the range of activations. For maximum efficiency, we adopt static quantization, with all the scaling factors fixed during inference.
Appendix B Error Term of Eq. 3
As one can see, the polynomial approximation of Eq. 3 exactly matches the data at the interpolating points . The error between a target function and the polynomial approximation is then:
where is some number that lies in the smallest interval containing . In general, this error reduces for large (for a properly selected set of interpolating points). Therefore, a sufficiently high-order polynomial that interpolates a target function is guaranteed to be a good approximation for it. We refer interested readers to for more details on polynomial interpolation.
Appendix C Experimental Details
In I-BERT, all the MatMul operations are performed with INT8 precision, and are accumulated to INT32 precision. Furthermore, the Embedding layer is kept at INT8 precision. Moreover, the non-linear operations (i.e., GELU, Softmax, and LayerNorm) are processed with INT32 precision, as we found that keeping them at high precision is important to ensure no accuracy degradation after quantization. Importantly, note that using INT32 for computing these operations has little overhead, as input data is already accumulated with INT32 precision, and these non-linear operations have linear computational complexity. We perform Requantization operation after these operations to bring the precision down from INT32 back to INT8 so that the follow up operations (e.g., next MatMuls) can be performed with low precision.
C.2 Training
C.3 Accuracy Evaluation on the GLUE Tasks
For evaluating the results, we use the standard metrics for each task in GLUE. In particular, we use classification accuracy and F1 score for QQP and MRPC , Pearson Correlation and Spearman Correlation for STS-B , and Mathews Correlation Coefficient for CoLA . For the remaining tasks , we use classification accuracy. For the tasks with multiple metrics, we report the average of them. Since there are two development sets for MNLI , i.e., MNLI-match (MNLI-m) for in-domain evaluation, and MNLI-mismatch (MNLI-mm) for cross-domain evaluation, and we report the accuracy on both datasets. We exclude WNLI as it has relatively small dataset and shows an unstable behaviour .
C.4 Environment Setup for Latency Evaluation
We use TensorRT 7.2.1 to deploy and tune the latency of BERT-Base and BERT-Large models (both INT8 and FP32) on Google Cloud Platform virtual machine with a single Tesla T4 GPU, CUDA 11.1, and cuDNN 8.0.
We should also mention that the most efficient way of implementing BERT with TensorRT is to use NVIDIA’s plugins that optimize and accelerate key operations in the Transformer architecture via operation fusion. Our estimates are that INT8 inference using NVIDIA’s plugins is about 2 times faster than naïvely using TensorRT APIs. However, we cannot modify those plugins to support our integer-only kernels as they are partially closed sourced and pre-compiled. Therefore, our latency evaluation is conducted without fully utilizing NVIDIA’s plugins. This leaves us a chance for further optimization to achieve our latency speedup relative to FP32 even more significant. As such, one could expect the potential for a further speed up with INT8 quantization.