Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference

Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, Dmitry Kalenichenko

Introduction

Current state-of-the-art Convolutional Neural Networks (CNNs) are not well suited for use on mobile devices. Since the advent of AlexNet , modern CNNs have primarily been appraised according to classification / detection accuracy. Thus network architectures have evolved without regard to model complexity and computational efficiency. On the other hand, successful deployment of CNNs on mobile platforms such as smartphones, AR/VR devices (HoloLens, Daydream), and drones require small model sizes to accommodate limited on-device memory, and low latency to maintain user engagement. This has led to a burgeoning field of research that focuses on reducing the model size and inference time of CNNs with minimal accuracy losses.

Approaches in this field roughly fall into two categories. The first category, exemplified by MobileNet , SqueezeNet , ShuffleNet , and DenseNet , designs novel network architectures that exploit computation / memory efficient operations. The second category quantizes the weights and / or activations of a CNN from 32 bit floating point into lower bit-depth representations. This methodology, embraced by approaches such as Ternary weight networks (TWN ), Binary Neural Networks (BNN ), XNOR-net , and more , is the focus of our investigation. Despite their abundance, current quantization approaches are lacking in two respects when it comes to trading off latency with accuracy.

First, prior approaches have not been evaluated on a reasonable baseline architecture. The most common baseline architectures, AlexNet , VGG and GoogleNet , are all over-parameterized by design in order to extract marginal accuracy improvements. Therefore, it is easy to obtain sizable compression of these architectures, reducing quantization experiments on these architectures to proof-of-concepts at best. Instead, a more meaningful challenge would be to quantize model architectures that are already efficient at trading off latency with accuracy, e.g. MobileNets.

Second, many quantization approaches do not deliver verifiable efficiency improvements on real hardware. Approaches that quantize only the weights () are primarily concerned with on-device storage and less with computational efficiency. Notable exceptions are binary, ternary and bit-shift networks . These latter approaches employ weights that are either 0 or powers of 2, which allow multiplication to be implemented by bit shifts. However, while bit-shifts can be efficient in custom hardware, they provide little benefit on existing hardware with multiply-add instructions that, when properly used (i.e. pipelined), are not more expensive than additions alone. Moreover, multiplications are only expensive if the operands are wide, and the need to avoid multiplications diminishes with bit depth once both weights and activations are quantized. Notably, these approaches rarely provide on-device measurements to verify the promised timing improvements. More runtime-friendly approaches quantize both the weights and the activations into 1 bit representations . With these approaches, both multiplications and additions can be implemented by efficient bit-shift and bit-count operations, which are showcased in custom GPU kernels (BNN ). However, 1 bit quantization often leads to substantial performance degradation, and may be overly stringent on model representation.

In this paper we address the above issues by improving the latency-vs-accuracy tradeoffs of MobileNets on common mobile hardware. Our specific contributions are:

We provide a quantization scheme (section 2.1) that quantizesh both weights and activations as 8-bit integers, and just a few parameters (bias vectors) as 32-bit integers.

We provide a quantized inference framework that is efficiently implementable on integer-arithmetic-only hardware such as the Qualcomm Hexagon (sections 2.2, 2.3), and we describe an efficient, accurate implementation on ARM NEON (Appendix B).

We provide a quantized training framework (section 3) co-designed with our quantized inference to minimize the loss of accuracy from quantization on real models.

We apply our frameworks to efficient classification and detection systems based on MobileNets and provide benchmark results on popular ARM CPUs (section 4) that show significant improvements in the latency-vs-accuracy tradeoffs for state-of-the-art MobileNet architectures, demonstrated in ImageNet classification , COCO object detection , and other tasks.

Our work draws inspiration from , which leverages low-precision fixed-point arithmetic to accelerate the training speed of CNNs, and from , which uses 88-bit fixed-point arithmetic to speed up inference on x86 CPUs. Our quantization scheme focuses instead on improving the inference speed vs accuracy tradeoff on mobile CPUs.

Quantized Inference

In this section, we describe our general quantization schemeThe quantization scheme described here is the one adopted in TensorFlow Lite and we will refer to specific parts of its code to illustrate aspects discussed below.We had earlier described this quantization scheme in the documentation of gemmlowp . That page may still be useful as an alternate treatment of some of the topics developed in this section, and for its self-contained example code., that is, the correspondence between the bit-representation of values (denoted qq below, for “quantized value”) and their interpretation as mathematical real numbers (denoted rr below, for “real value”). Our quantization scheme is implemented using integer-only arithmetic during inference and floating-point arithmetic during training, with both implementations maintaining a high degree of correspondence with each other. We achieve this by first providing a mathematically rigorous definition of our quantization scheme, and separately adopting this scheme for both integer-arithmetic inference and floating-point training.

A basic requirement of our quantization scheme is that it permits efficient implementation of all arithmetic using only integer arithmetic operations on the quantized values (we eschew implementations requiring lookup tables because these tend to perform poorly compared to pure arithmetic on SIMD hardware). This is equivalent to requiring that the quantization scheme be an affine mapping of integers qq to real numbers rr, i.e. of the form

for some constants SS and ZZ. Equation (1) is our quantization scheme and the constants SS and ZZ are our quantization parameters. Our quantization scheme uses a single set of quantization parameters for all values within each activations array and within each weights array; separate arrays use separate quantization parameters.

For 8-bit quantization, qq is quantized as an 8-bit integer (for BB-bit quantization, qq is quantized as an BB-bit integer). Some arrays, typically bias vectors, are quantized as 32-bit integers, see section 2.4.

The constant SS (for “scale”) is an arbitrary positive real number. It is typically represented in software as a floating-point quantity, like the real values rr. Section 2.2 describes methods for avoiding the representation of such floating-point quantities in the inference workload.

The constant ZZ (for “zero-point”) is of the same type as quantized values qq, and is in fact the quantized value qq corresponding to the real value 0. This allows us to automatically meet the requirement that the real value r=0r=0 be exactly representable by a quantized value. The motivation for this requirement is that efficient implementation of neural network operators often requires zero-padding of arrays around boundaries.

Our discussion so far is summarized in the following quantized buffer data structureThe actual data structures in the TensorFlow Lite Converter are QuantizationParams and Array in this header file. As we discuss in the next subsection, this data structure, which still contains a floating-point quantity, does not appear in the actual quantized on-device inference code., with one instance of such a buffer existing for each activations array and weights array in a neural network. We use C++ syntax because it allows the unambiguous conveyance of types.

template // e.g. QType=uint8struct QuantizedBuffer { vector q; // the quantized values float S; // the scale QType Z; // the zero-point};

2 Integer-arithmetic-only matrix multiplication

We now turn to the question of how to perform inference using only integer arithmetic, i.e. how to use Equation (1) to translate real-numbers computation into quantized-values computation, and how the latter can be designed to involve only integer arithmetic even though the scale values SS are not integers.

Consider the multiplication of two square N×NN\times N matrices of real numbers, r1r_{1} and r2r_{2}, with their product represented by r3=r1r2r_{3}=r_{1}r_{2}. We denote the entries of each of these matrices rαr_{\alpha} (α=1\alpha=1, 22 or 33) as rα(i,j)r_{\alpha}^{(i,j)} for 1⩽i,j⩽N1\leqslant i,j\leqslant N, and the quantization parameters with which they are quantized as (Sα,Zα)(S_{\alpha},Z_{\alpha}). We denote the quantized entries by qα(i,j)q_{\alpha}^{(i,j)}. Equation (1) then becomes:

From the definition of matrix multiplication, we have

In Equation (4), the only non-integer is the multiplier MM. As a constant depending only on the quantization scales S1,S2,S3S_{1},S_{2},S_{3}, it can be computed offline. We empirically find it to always be in the interval (0,1)(0,1), and can therefore express it in the normalized form

where M0M_{0} is in the interval [0.5,1)[0.5,1) and nn is a non-negative integer. The normalized multiplier M0M_{0} now lends itself well to being expressed as a fixed-point multiplier (e.g. int16 or int32 depending on hardware capability). For example, if int32 is used, the integer representing M0M_{0} is the int32 value nearest to 231M02^{31}M_{0}. Since M0⩾0.5M_{0}\geqslant 0.5, this value is always at least 2302^{30} and will therefore always have at least 30 bits of relative accuracy. Multiplication by M0M_{0} can thus be implemented as a fixed-point multiplicationThe computation discussed in this section is implemented in TensorFlow Lite reference code for a fully-connected layer.. Meanwhile, multiplication by 2−n2^{-n} can be implemented with an efficient bit-shift, albeit one that needs to have correct round-to-nearest behavior, an issue that we return to in Appendix B.

3 Efficient handling of zero-points

In order to efficiently implement the evaluation of Equation (4) without having to perform 2N32N^{3} subtractions and without having to expand the operands of the multiplication into 16-bit integers, we first notice that by distributing the multiplication in Equation (4), we can rewrite it as

Each a2(k)a_{2}^{(k)} or aˉ1(i)\bar{a}_{1}^{(i)} takes only NN additions to compute, so they collectively take only 2N22N^{2} additions. The rest of the cost of the evaluation of (7) is almost entirely concentrated in the core integer matrix multiplication accumulation

which takes 2N32N^{3} arithmetic operations; indeed, everything else involved in (7) is O(N2)O(N^{2}) with a small constant in the OO. Thus, the expansion into the form (7) and the factored-out computation of a2(k)a_{2}^{(k)} and aˉ1(i)\bar{a}_{1}^{(i)} enable low-overhead handling of arbitrary zero-points for anything but the smallest values of NN, reducing the problem to the same core integer matrix multiplication accumulation (9) as we would have to compute in any other zero-points-free quantization scheme.

4 Implementation of a typical fused layer

We continue the discussion of section 2.3, but now explicitly define the data types of all quantities involved, and modify the quantized matrix multiplication (7) to merge the bias-addition and activation function evaluation directly into it. This fusing of whole layers into a single operation is not only an optimization. As we must reproduce in inference code the same arithmetic that is used in training, the granularity of fused operators in inference code (taking an 8-bit quantized input and producing an 8-bit quantized output) must match the placement of “fake quantization” operators in the training graph (section 3).

For our implementation on ARM and x86 CPU architectures, we use the gemmlowp library , whose GemmWithOutputPipeline entry point provides supports the fused operations that we now describeThe discussion in this section is implemented in TensorFlow Lite for e.g. a Convolutional operator (reference code is self-contained, optimized code calls into gemmlowp )..

We take the q1q_{1} matrix to be the weights, and the q2q_{2} matrix to be the activations. Both the weights and activations are of type uint8 (we could have equivalently chosen int8, with suitably modified zero-points). Accumulating products of uint8 values requires a 32-bit accumulator, and we choose a signed type for the accumulator for a reason that will soon become clear. The sum in (9) is thus of the form:

Although the bias-vectors are quantized as 32-bit values, they account for only a tiny fraction of the parameters in a neural network. Furthermore, the use of higher precision for bias vectors meets a real need: as each bias-vector entry is added to many output activations, any quantization error in the bias-vector tends to act as an overall bias (i.e. an error term with nonzero mean), which must be avoided in order to preserve good end-to-end neural network accuracyThe quantization of bias-vectors discussed here is implemented here in the TensorFlow Lite Converter..

With the final value of the int32 accumulator, there remain three things left to do: scale down to the final scale used by the 8-bit output activations, cast down to uint8 and apply the activation function to yield the final 8-bit output activation.

The down-scaling corresponds to multiplication by the multiplier MM in equation (7). As explained in section 2.2, it is implemented as a fixed-point multiplication by a normalized multiplier M0M_{0} and a rounding bit-shift. Afterwards, we perform a saturating cast to uint8, saturating to the range $$.

We focus on activation functions that are mere clamps, e.g. ReLU, ReLU6. Mathematical functions are discussed in appendix A.1 and we do not currently fuse them into such layers. Thus, the only thing that our fused activation functions need to do is to further clamp the uint8 value to some sub-interval of beforestoringthefinaluint8outputactivation.Inpractice,thequantizedtrainingprocess(section3)tendstolearntomakeuseofthewholeoutputuint8before storing the final uint8 output activation. In practice, the quantized training process (section 3) tends to learn to make use of the whole output uint8 interval so that the activation function no longer does anything, its effect being subsumed in the clamping to $$ implied in the saturating cast to uint8.

Training with simulated quantization

A common approach to training quantized networks is to train in floating point and then quantize the resulting weights (sometimes with additional post-quantization training for fine-tuning). We found that this approach works sufficiently well for large models with considerable representational capacity, but leads to significant accuracy drops for small models. Common failure modes for simple post-training quantization include: 1) large differences (more than 100×100\times) in ranges of weights for different output channels (section 2 mandates that all channels of the same layer be quantized to the same resolution, which causes weights in channels with smaller ranges to have much higher relative error) and 2) outlier weight values that make all remaining weights less precise after quantization.

We propose an approach that simulates quantization effects in the forward pass of training. Backpropagation still happens as usual, and all weights and biases are stored in floating point so that they can be easily nudged by small amounts. The forward propagation pass however simulates quantized inference as it will happen in the inference engine, by implementing in floating-point arithmetic the rounding behavior of the quantization scheme that we introduced in section 2:

Weights are quantized before they are convolved with the input. If batch normalization (see ) is used for the layer, the batch normalization parameters are “folded into” the weights before quantization, see section 3.2.

Activations are quantized at points where they would be during inference, e.g. after the activation function is applied to a convolutional or fully connected layer’s output, or after a bypass connection adds or concatenates the outputs of several layers together such as in ResNets.

For each layer, quantization is parameterized by the number of quantization levels and clamping range, and is performed by applying point-wise the quantization function qq defined as follows:

where rr is a real-valued number to be quantized, [a;b][a;b] is the quantization range, nn is the number of quantization levels, and ⌊⋅⌉\lfloor\cdot\rceil denotes rounding to the nearest integer. nn is fixed for all layers in our experiments, e.g. n=28=256n=2^{8}=256 for 8 bit quantization.

Quantization ranges are treated differently for weight quantization vs. activation quantization:

For weights, the basic idea is simply to set a≔min⁡wa\coloneqq\min w, b≔max⁡wb\coloneqq\max w. We apply a minor tweak to this so that the weights, once quantized as int8 values, only range in $andnevertakethevalueand never take the value-128$, as this enables a substantial optimization opportunity (for more details, see Appendix B).

For activations, ranges depend on the inputs to the network. To estimate the ranges, we collect [a;b][a;b] ranges seen on activations during training and then aggregate them via exponential moving averages (EMA) with the smoothing parameter being close to 1 so that observed ranges are smoothed across thousands of training steps. Given significant delay in the EMA updating activation ranges when the ranges shift rapidly, we found it useful to completely disable activation quantization at the start of training (say, for 50 thousand to 2 million steps). This allows the network to enter a more stable state where activation quantization ranges do not exclude a significant fraction of values.

In both cases, the boundaries [a;b][a;b] are nudged so that value 0.00.0 is exactly representable as an integer z(a,b,n)z(a,b,n) after quantization. As a result, the learned quantization parameters map to the scale SS and zero-point ZZ in equation 1:

Below we depict simulated quantization assuming that the computations of a neural network are captured as a TensorFlow graph . A typical workflow is described in Algorithm 1.

Optimization of the inference graph by fusing and removing operations is outside the scope of this paper. Source code for graph modifications (inserting fake quantization operations, creating and optimizing the inference graph) and a low bit inference engine has been open-sourced with TensorFlow contributions in .

Figure 1.1a and b illustrate TensorFlow graphs before and after quantization for a simple convolutional layer. Illustrations of the more complex convolution with a bypass connection in figure C.3 can be found in figure C.4.

Note that the biases are not quantized because they are represented as 32-bit integers in the inference process, with a much higher range and precision compared to the 8 bit weights and activations. Furthermore, quantization parameters used for biases are inferred from the quantization parameters of the weights and activations. See section 2.4.

Typical TensorFlow code illustrating use of follows:

2 Batch normalization folding

For models that use batch normalization (see ), there is additional complexity: the training graph contains batch normalization as a separate block of operations, whereas the inference graph has batch normalization parameters “folded” into the convolutional or fully connected layer’s weights and biases, for efficiency. To accurately simulate quantization effects, we need to simulate this folding, and quantize weights after they have been scaled by the batch normalization parameters. We do so with the following:

Here γ\gamma is the batch normalization’s scale parameter, EMA(σB2)EMA(\sigma^{2}_{B}) is the moving average estimate of the variance of convolution results across the batch, and ε\varepsilon is just a small constant for numerical stability.

Experiments

We conducted two set of experiments, one showcasing the effectiveness of quantized training (Section. 4.1), and the other illustrating the improved latency-vs-accuracy tradeoff of quantized models on common hardware (Section. 4.2). The most performance-critical part of the inference workload on the neural networks being benchmarked is matrix multiplication (GEMM). The 8-bit and 32-bit floating-point GEMM inference code uses the gemmlowp library for 8-bit quantized inference, and the Eigen library for 32-bit floating-point inference.

We apply quantized training to ResNets and InceptionV3 on the ImageNet dataset. These popular networks are too computationally intensive to be deployed on mobile devices, but are included for comparison purposes. Training protocols are discussed in Appendix D.1 and D.2.

We compare floating-point vs integer-quantized ResNets for various depths in table 4.1. Accuracies of integer-only quantized networks are within 2%2\% of their floating-point counterparts.

We also list ResNet50 accuracies under different quantization schemes in table 4.2. As expected, integer-only quantization outperforms FGQ , which uses 2 bits for weight quantization. INQ (5-bit weight floating-point activation) achieves a similar accuracy as ours, but we provide additional run-time improvements (see section 4.2).

1.2 Inception v3 on ImageNet

We compare the Inception v3 model quantized into 8 and 7 bits, respectively. 7-bit quantization is obtained by setting the number of quantization levels in equation 12 to n=27n=2^{7}. We additionally probe the sensitivity of activation quantization by comparing networks with two activation nonlinearities, ReLU6 and ReLU. The training protocol is in Appendix D.2.

Table 4.3 shows that 7-bit quantized training produces model accuracies close to that of 8-bit quantized training, and quantized models with ReLU6 have less accuracy degradation. The latter can be explained by noticing that ReLU6 introduces the interval $$ as a natural range for activations, while ReLU allows activations to take values from a possibly larger interval, with different ranges in different channels. Values in a fixed range are easier to quantize with high precision.

2 Quantization of MobileNets

MobileNets are a family of architectures that achieve a state-of-the-art tradeoff between on-device latency and ImageNet classification accuracy. In this section we demonstrate how integer-only quantization can further improve the tradeoff on common hardware.

We benchmarked the MobileNet architecture with varying depth-multipliers (DM) and resolutions on ImageNet on three types of Qualcomm cores, which represent three different micro-architectures: 1) Snapdragon 835 LITTLE core, (figure. 1.1c), a power-efficient processor found in Google Pixel 2; 2) Snapdragon 835 big core (figure. 4.1), a high-performance core employed by Google Pixel 2; and 3) Snapdragon 821 big core (figure. 4.2), a high-performance core used in Google Pixel 1.

Integer-only quantized MobileNets achieve higher accuracies than floating-point MobileNets given the same runtime budget. The accuracy gap is quite substantial (∼10%\sim 10\%) for Snapdragon 835 LITTLE cores at the 33ms latency needed for real-time (30 fps) operation. While most of the quantization literature focuses on minimizing accuracy loss for a given architecture, we advocate for a more comprehensive latency-vs-accuracy tradeoff as a better measure. Note that this tradeoff depends critically on the relative speed of floating-point vs integer-only arithmetic in hardware. Floating-point computation is better optimized in the Snapdragon 821, for example, resulting in a less noticeable reduction in latency for quantized models.

2.2 COCO

We evaluated quantization in the context of mobile real time object detection, comparing the performance of quantized 8-bit and float models of MobileNet SSD on the COCO dataset . We replaced all the regular convolutions in the SSD prediction layers with separable convolutions (depthwise followed by 1×11\times 1 projection). This modification is consistent with the overall design of MobileNets and makes them more computationally efficient. We utilized the Open Source TensorFlow Object Detection API to train and evaluate our models. The training protocol is described in Appendix D.3. We also delayed quantization for 500500 thousand steps (see section 3.1), finding that it significantly decreases the time to convergence.

Table 4.4 shows the latency-vs-accuracy tradeoff between floating-point and integer-quantized models. Latency was measured on a single thread using Snapdragon 835 cores (big and LITTLE). Quantized training and inference results in up to a 50%50\% reduction in running time, with a minimal loss in accuracy (−1.8%-1.8\% relative).

2.3 Face detection

To better examine quantized MobileNet SSD on a smaller scale, we benchmarked face detection on the face attribute classification dataset (a Flickr-based dataset used in ). We contacted the authors of to evaluate our quantized MobileNets on detection and face attributes following the same protocols (detailed in Appendix D.4).

As indicated by tables 4.5 and 4.6, quantization provides close to a 2×2\times latency reduction with a Qualcomm Snapdragon 835 big or LITTLE core at the cost of a ∼2%\sim 2\% drop in the average precision. Notably, quantization allows the 25%25\% face detector to run in real-time (1K/28≈361K/28\approx 36 fps) on a single big core, whereas the floating-point model remains slower than real-time (1K/44≈231K/44\approx 23 fps).

We additionally examine the effect of multi-threading on the latency of quantized models. Table 4.6 shows a 1.51.5 to 2.2×2.2\times) speedup when using 44 cores. The speedup ratios are comparable between the two cores, and are higher for larger models where the overhead of multi-threading occupies a smaller fraction of the total computation.

2.4 Face attributes

Figure 4.3 shows the latency-vs-accuracy tradeoff of face attribute classification on the Qualcomm Snapdragon 821. Since quantized training results in little accuracy degradation, we see an improved tradeoff even though the Qualcomm Snapdragon 821 is highly optimized for floating point arithmetic (see Figure 4.2 for comparison).

Ablation study To understand performance sensitivity to the quantization scheme, we further evaluate quantized training with varying weight and activation quantization bit depths. The degradation in average precision for binary attributes and age precision relative to the floating-point baseline are shown in Tables 4.7 and 4.8, respectively. The tables suggest that 1) weights are more sensitive to reduced quantization bit depth than activations, 2) 8 and 7-bit quantized models perform similarly to floating point models, and 3) when the total bit-depths are equal, it is better to keep weight and activation bit depths the same.

Discussion

We propose a quantization scheme that relies only on integer arithmetic to approximate the floating-point computations in a neural network. Training that simulates the effect of quantization helps to restore model accuracy to near-identical levels as the original. In addition to the 4×4\times reduction of model size, inference efficiency is improved via ARM NEON-based implementations. The improvement advances the state-of-the-art tradeoff between latency on common ARM CPUs and the accuracy of popular computer vision models. The synergy between our quantization scheme and efficient architecture design suggests that integer-arithmetic-only inference could be a key enabler that propels visual recognition technologies into the real-time and low-end phone market.

References

Appendix A Appendix: Layer-specific details

Math functions such as hyperbolic tangent, the logistic function, and softmax often appear in neural networks. No lookup tables are needed since these functions are implemented in pure fixed-point arithmetic similarly to how they would be implemented in floating-point arithmeticPure-arithmetic, SIMD-ready, branch-free, fixed-point implementations of at least tanh and the logistic functions are given in gemmlowp ’s fixedpoint directory, with specializations for NEON and SSE instruction sets. One can see in TensorFlow Lite how these are called..

A.2 Addition

Some neural networks use a plain Addition layer type, that simply adds two activation arrays together. Such Addition layers are more expensive in quantized inference compared to floating-point because rescaling is needed: one input needs to be rescaled onto the other’s scale using a fixed-point multiplication by the multiplier M=S1/S2M=S_{1}/S_{2} similar to what we have seen earlier (end of section 2.2), before the actual addition can be performed as a simple integer addition; finally, the result must be rescaled again to fit the output array’s scaleSee the TensorFlow Lite implementation..

A.3 Concatenation

Fully general support for concatenation layers poses the same rescaling problem as Addition layers. Because such rescaling of uint8 values would be a lossy operation, and as it seems that concatenation ought to be a lossless operation, we prefer to handle this problem differently: instead of implementing lossy rescaling, we introduce a requirement that all the input activations and the output activations in a Concatenation layer have the same quantization parameters. This removes the need for rescaling and concatenations are thus lossless and free of any arithmeticThis is implemented in this part of the TensorFlow Lite Converter.

Appendix B Appendix: ARM NEON details

This section assumes familiarity with assembly programming on the ARM NEON instruction set. The instruction mnemonics below refer to the 64-bit ARM instruction set, but the discussion applies equally to 32-bit ARM instructions.

The fixed-point multiplications referenced throughout this article map exactly to the SQRDMULH instruction. It is very important to use the correctly-rounding instruction SQRDMULH and not SQDMULHThe fixed-point math function implementations in gemmlowp use such fixed-point multiplications, and ordinary (non-saturating) integer additions. We have no use for general saturated arithmetic..

The rounding-to-nearest right-shifts referenced in section 2.2 do not map exactly to any ARM NEON instruction. The problem is that the “rounding right shift” instruction, RSHL with variable negative offset, breaks ties by rounding upward, instead of rounding them away from zero. For example, if we use RSHL to implement the division −12/23-12/2^{3}, the result will be −1-1 whereas it should be −2-2 with “round to nearest”. This is problematic as it results in an overall upward bias, which has been observed to cause significant loss of end-to-end accuracy in neural network inference. A correct round-to-nearest right-shift can still be implemented using RSHL but with suitable fix-up arithmetic around itIt is implemented here in gemmlowp ..

For efficient NEON implementation of the matrix multiplication’s core accumulation, we use the following trick. In the multiply-add operation in (10), we first change the operands’ type from uint8 to int8 (which can be done by subtracting 128 from the quantized values and zero-points). Thus the core multiply-add becomes

As mentioned in section 3, with a minor tweak of the quantized training process, we can ensure that the weights, once quantized as int8 values, never take the value −128-128. Hence, the product in (B.1) is never −128∗−128-128*-128, and is therefore always less than 2142^{14} in absolute value. Hence, (B.1) can accumulate two products on a local int16 accumulator before that needs to be accumulated into the true int32 accumulator. This allows the use of an 8-way SIMD multiplication (SMULL on int8 operands), followed by an 8-way SIMD multiply-add (SMLAL on int8 operands), followed by a pairwise-add-and-accumulate into the int32 accumulators (SADALP)This technique is implemented in the optimized NEON kernel in gemmlowp , which is in particular what TensorFlow Lite uses (see the choice of L8R8WithLhsNonzeroBitDepthParams at this line)..

Appendix C Appendix: Graph diagrams

Appendix D Experimental protocols

Preprocessing. All images from ImageNet are resized preserving aspect ratio so that the smallest side of the image is 256256. Then the center 224×224224\times 224 patch is cropped and the means are subtracted for each of the RGB channels.

Optimization. We use the momentum optimizer from TensorFlow with momentum 0.90.9 and a batch size of 3232. The learning rate starts from 10−510^{-5} and decays in a staircase fashion by 0.10.1 for every 3030 epochs. Activation quantization is delayed for 500,000500,000 steps for reasons discussed in section 3. Training uses 5050 workers asynchronously, and stops after validation accuracy plateaus, normally after 100100 epochs.

D.2 Inception protocol

All results in table 4.3 were obtained after training for approximately 1010 million steps, with batches of 3232 samples, using 5050 distributed workers, asynchronously. Training data were ImageNet 2012 299×299299\times 299 images with labels. Image augmentation consisted of: random crops, random horizontal flips, and random color distortion. The optimizer used was RMSProp with learning rate starting at 0.0450.045 and decaying exponentially and stepwise with factor 0.940.94 after every 22 epochs. Other RMSProp parameters were: 0.90.9 momentum, 0.90.9 decay, 1.01.0 epsilon term. Trained parameters were EMA averaged with decay 0.99990.9999.

D.3 COCO detection protocol

Preprocessing. During training, all images are randomly cropped and resized to 320×320320\times 320. During evaluation, all images are directly resized to 320×320320\times 320. All input values are normalized to $$.

Optimization. We used the RMSprop optimizer from TensorFlow with a batch size of 3232. The learning rate starts from 4×10−34\times 10^{-3} and decays in a staircase fashion by a factor of 0.10.1 for every 100100 epochs. Activation quantization is delayed for 500,000500,000 steps for reasons discussed in section 3. Training uses 2020 workers asynchronously, and stops after validation accuracy plateaus, normally after approximately 66 million steps.

Metrics. Evaluation results are reported with the COCO primary challenge metric: AP at IoU=.50:.05:.95. We follow the same train/eval split in .

D.4 Face detection and face attribute classification protocol

Preprocessing. Random 1:1 crops are taken from images in the Flickr-based dataset used in and resized to 320×320320\times 320 pixels for face detection and 128×128128\times 128 pixels for face attribute classification. The resulting crops are flipped horizontally with a 50%50\% probability. The values for each of the RGB channels are renormalized to be in the range $$.

Face Detection Optimization. We used the RMSprop optimizer from TensorFlow with a batch size of 3232. The learning rate starts from 4×10−34\times 10^{-3} and decays in a staircase fashion by a factor of 0.10.1 for every 100100 epochs. Activation quantization is delayed for 500,000500,000 steps for reasons discussed in section 3. Training uses 2020 workers asynchronously, and stops after validation accuracy plateaus, normally after approximately 33 million steps.

Face Attribute Classification Optimization. We followed the optimization protocol in . We used the Adagrad optimizer from Tensorflow with a batch size of 3232 and a constant learning rate of 0.10.1. Training uses 1212 workers asynchronously, and stops at 2020 million steps.

Latency Measurements. We created a binary that runs the face detection and face attributes classification models repeatedly on random inputs for 100100 seconds. We pushed this binary to Pixel and Pixel 2 phones using the adb push command, and executed it on 1, 2, and 4 LITTLE cores, and 1, 2, and 4 big cores using the adb shell command with the appropriate taskset specified. We reported the average runtime of the face detector model on 320×320320\times 320 inputs, and of the face attributes classifier model on 128×128128\times 128 inputs.