Joint Autoregressive and Hierarchical Priors for Learned Image Compression
David Minnen, Johannes Ballé, George Toderici
Introduction
Most recent methods for learning-based, lossy image compression adopt an approach based on transform coding . In this approach, image compression is achieved by first mapping pixel data into a quantized latent representation and then losslessly compressing the latents. Within the deep learning research community, the transforms typically take the form of convolutional neural networks (CNNs), which approximate nonlinear functions with the potential to map pixels into a more compressible latent space than the linear transforms used by traditional image codecs. This nonlinear transform coding method resembles an autoencoder , which consists of an encoder transform between the data (in this case, pixels) and a latent, reduced-dimensionality space, and a decoder, an approximate inverse function that maps latents back to pixels. While dimensionality reduction can be seen as a simplistic form of compression, it is not equivalent to it, as the goal of compression is to reduce the entropy of the representation under a prior probability model shared between the sender and the receiver (the entropy model), not only the dimensionality. To improve compression performance, recent methods have given increased focus to this part of the model . Finally, the entropy model is used in conjunction with standard entropy coding algorithms such as arithmetic, range, or Huffman coding to generate a compressed bitstream.
The training goal is to minimize the expected length of the bitstream as well as the expected distortion of the reconstructed image with respect to the original, giving rise to a rate–distortion optimization problem:
where is the Lagrange multiplier that determines the desired rate–distortion trade-off, is the unknown distribution of natural images, represents rounding to the nearest integer (quantization), is the encoder, are the quantized latents, is a discrete entropy model, and is the decoder with representing the reconstructed image. The rate term corresponds to the cross entropy between the marginal distribution of the latents and the learned entropy model, which is minimized when the two distributions are identical. The distortion term may correspond to a closed-form likelihood, such as when represents mean squared error (MSE), which induces an interpretation of the model as a variational autoencoder . When optimizing the model for other distortion metrics such as MS-SSIM, it is simply minimized as an energy function.
The models we analyze in this paper build on the work of Ballé et al. , which uses a noise-based relaxation to be able to apply gradient descent methods to the loss function in Eq. (1) and introduces a hierarchical prior to improve the entropy model. While most previous research uses a fixed, though potentially complex, entropy model, Ballé et al. use a Gaussian scale mixture (GSM) where the scale parameters are conditioned on a hyperprior. Their model allows for end-to-end training, which includes joint optimization of a quantized representation of the hyperprior, the conditional entropy model, and the base autoencoder. The key insight is that the compressed hyperprior could be added to the generated bitstream as side information, which allows the decoder to use the conditional entropy model. In this way, the entropy model itself is image-dependent and spatially adaptive, which allows for a richer and more accurate model. Ballé et al. show that standard optimization methods for deep neural networks are sufficient to learn a useful balance between the size of the side information and the savings gained from a more accurate entropy model. The resulting compression model provides state-of-the-art image compression results compared to earlier learning-based methods.
We extend this GSM-based entropy model in two ways: first, by generalizing the hierarchical GSM model to a Gaussian mixture model, and, inspired by recent work on generative models, by adding an autoregressive component. We assess the compression performance of both approaches, including variations in the network architectures, and discuss benefits and potential drawbacks of both extensions. For the results in this paper, we did not make efforts to reduce the capacity (i.e., number of channels, layers) of the artificial neural networks to optimize computational complexity, since we are interested in determining the potential of different forms of priors rather than trading off complexity against performance. Note that increasing capacity alone is not sufficient to obtain arbitrarily good compression performance [13, appendix 6.3].
Architecture Details
Figure 1 provides a high-level overview of our generalized compression model, which contains two main sub-networksSee Section 4 in the supplemental materials for an in-depth visual comparison between our architecture variants and previous learning-based methods.. The first is the core autoencoder, which learns a quantized latent representation of images (Encoder and Decoder blocks). The second sub-network is responsible for learning a probabilistic model over quantized latents used for entropy coding. It combines the Context Model, an autoregressive model over latents, with the hyper-network (Hyper Encoder and Hyper Decoder blocks), which learns to represent information useful for correcting the context-based predictions. The data from these two sources is combined by the Entropy Parameters network, which generates the mean and scale parameters for a conditional Gaussian entropy model.
Once training is complete, a valid compression model must prevent any information from passing between the encoder to the decoder unless that information is available in the compressed file. In Figure 1, the arithmetic encoding (AE) blocks produce the compressed representation of the symbols coming from the quantizer, which is stored in a file. Therefore at decoding time, any information that depends on the quantized latents may be used by the decoder once it has been decoded. In order for the context model to work, at any point it can only access the latents that have already been decoded. When starting to decode an image, we assume that the previously decoded latents have all been set to zero.
The learning problem is to minimize the expected rate–distortion loss defined in Eq. 1 over the model parameters. Following the work of Ballé et al. , we model each latent, , as a Gaussian convolved with a unit uniform distribution. This ensures a good match between encoder and decoder distributions of both the quantized latents, and continuous-valued latents subjected to additive uniform noise during training. While predicted the scale of each Gaussian conditioned on the hyperprior, , we extend the model by predicting the mean and scale parameters conditioned on both the hyperprior as well as the causal context of each latent , which we denote . The predicted Gaussian parameters are functions of the learned parameters of the hyper-decoder, context model, and entropy parameters networks (, , and , respectively):
The entropy model for the hyperprior is the same as in , although we expect the hyper-encoder and hyper-decoder to learn significantly different functions in our combined model, since they now work in conjunction with an autoregressive network to predict the entropy model parameters. Since we do not make any assumptions about the distribution of the hyper-latents, a non-parametric, fully factorized density model is used. A more powerful entropy model for the hyper-latents may improve compression rates, e.g., we could stack multiple instances of our contextual model, but we expect the net effect to be minimal since comprises only a very small percentage of the total file size. Because both the compressed latents and the compressed hyper-latents are part of the generated bitstream, the rate–distortion loss from Equation 1 must be expanded to include the cost of transmitting . Coupled with a squared error distortion metric, the full loss function becomes:
Details about the individual network layers in each component of our models are outlined in Table 1. While the internal structure of the components is fairly unrestricted, e.g., one could exchange the convolutional layers for residual blocks or dilated convolution without fundamentally changing the model, certain components must be constrained to ensure that availability of the bitstreams alone is sufficient for the receiver to reconstruct the image.
The last layer of the encoder corresponds to the bottleneck of the base autoencoder. Its number of output channels determines the number of elements that must be compressed and stored. Depending on the rate–distortion trade-off, our models learn to ignore certain channels by deterministically generating the same latent value and assigning it a probability of 1, which wastes computation but requires no additional entropy. This modeling flexibility allows us to set the bottleneck larger than necessary, and let the model determine the number of channels that yield the best performance. Similar to reported in other work, we found that too few channels in the bottleneck can impede rate–distortion performance when training models to target higher bit rates, but too having too many does not harm the compression performance.
The final layer of the decoder must have three channels to generate RGB images, and the final layer of the Entropy Parameters sub-network must have exactly twice as many channels as the bottleneck. This constraint arises because the Entropy Parameters network predicts two values, the mean and scale of a Gaussian distribution, for each latent. The number of output channels of the Context Model and Hyper Decoder components are not constrained, but we also set them to twice the bottleneck size in all of our experiments.
Although the formal definition of our model allows the autoregressive component to condition its predictions on all previous latents, in practice we use a limited context (55 convolution kernels) with masked convolution similar to the approach used by PixelCNN . The Entropy Parameters network is also constrained, since it can not access predictions from the Context Model beyond the current latent element. For simplicity, we use 11 convolution in the Entropy Parameters network, although masked convolution is also permissible. Section 3 provides an empirical evaluation of the model variants we assessed, exploring the effects of different context sizes and more complex autoregressive networks.
Experimental Results
We evaluate our generalized models by calculating the rate–distortion (RD) performance averaged over the publicly available Kodak image set Please see the supplemental material for additional evaluation results including full-page RD curves, example images, and results on the larger Tecnick image set (100 images with resolution 12001200).. Figure 2 shows RD curves using peak signal-to-noise ratio (PSNR) as the image quality metric. While PSNR is known to be a relatively poor perceptual metric , it is still a standard metric used to evaluate image compression algorithms and is the primary metric used for tuning conventional compression methods. The RD graph on the left of Figure 2 compares our combined context + hyperprior model to existing image codecs (standard codecs and learned models) and shows that this model outperforms all of the existing methods including BPG , a state-of-the-art codec based on the intra-frame coding algorithm from HEVC [HEVC]. To the best of our knowledge, it is the first learning-based compression model to outperform BPG on PSNR. The right RD graph compares different versions of our models and shows that the combined model performs the best, while the context-only model performs slightly worse than either hierachical version.
Figure 4 shows RD curves for Kodak using multiscale structural similarity (MS-SSIM) as the image quality metric. The graph includes two versions of our combined model: one optimized for MSE and one optimized for MS-SSIM. The latter outperforms all existing methods including all standard codecs and other learning-based methods that were also optimized for MS-SSIM (). As expected, when our model is optimized for MSE, performance according to MS-SSIM falls. Nonetheless, the MS-SSIM scores for this model still exceed all standard codecs and all learning-based methods that were not specifically optimized for MS-SSIM.
As outlined in Table 1, our baseline architecture for the combined model uses 55 masked convolution in a single linear layer for the context model, and it uses a conditional Gaussian distribution for the entropy model. Figure 4 compares this baseline to several variants by showing the relative increase in file size at a single rate-point. The green bars show that exchanging the Gaussian distribution for a logistic distribution has almost no effect (the 0.3% increase is smaller than the training variance), while switching to a Laplacian distribution decreases performance more substantially. The blue bars compare different context configurations. Masked 33 and 77 convolution both perform slightly worse, which is surprising since we expected the additional context provided by the 77 kernels to improve prediction accuracy. Similarly, a 3-layer, nonlinear context model using 55 masked convolution also performed slightly worse than the linear baseline. Finally, the purple bars show the effect of using a severely restricted context such as only a single neighbor or three neighbors from the previous row. The primary benefit of these models is increased parallelization when calculating context-based predictions since the dependence is reduced from two dimensions down to one. While both cases show a non-negligible rate increase (2.1% and 3.1%, respectively), the increase may be worthwhile in a practical implementation where runtime speed is a major concern.
Finally, Figure 5 provides a visual comparison for one of the Kodak images. Creating accurate comparisons is difficult since most compression methods do not have the ability to target a precise bit rate. We therefore selected comparison images with sizes that are as close as possible, but always larger than our encoding (up to 9.4% larger in the case of BPG). Nonetheless, our compression model provides clearly better visual quality compared to the scale hyperprior baseline and JPEG. The perceptual quality relative to BPG is much closer. For example, BPG preserves mode detail in the sky and parts of the fence, but at the expense of introducing geometric artifacts in the sky, mild ringing near the building/sky boundaries, and some boundary artifacts where neighboring blocks have widely different levels of detail (e.g., in the grass and lighthouse).
Related Work
The earliest research that used neural networks to compress images dates back to the 1980s and relies on an autoencoder with a small bottleneck using either uniform quantization or vector quantization . These approaches sought equal utilization of the codes and thus did not learn an explicit entropy model. Considerable research followed these initial models, and Jiang provides a comprehensive survey covering methods published through the late 1990s .
More recently, image compression with deep neural networks became a popular research topic starting with the work of Toderici et al. who used a recurrent architecture based on LSTMs to learn multi-rate, progressive models. Their approach was improved by exploring other recurrent architectures for the autoencoder, training an LSTM-based entropy model, and adding a post-process that spatially adapts the bit rate based on the complexity of the local image content . Related research followed a more traditional image coding approach and explicitly divided images into patches instead of using a fully convolutional model . Inspired by modern image codecs and learned inpainting algorithms, these methods trained a neural network to predict each image patch from its causal context (in the image space, not the latent space) before encoding the residual. Similarly, most modern image compression standards use context to predict pixel values as well as using a context-adaptive entropy model .
Many learning-based methods take the form of an autoencoder, and multiple models are trained to target different bit rates instead of training a single recurrent model . Some use a fully factorized entropy model , while others make use of context in code space to improve compression rates . Other methods do not make use of context via an autoregressive model and instead rely on side information that is either predicted by a neural network or composed of indices into a (shared) dictionary of non-parametric code distributions used locally by the entropy coder .
A further major difference are the constraints imposed on compression models by the need to quantize and arithmetically encode the latents, which require certain choices regarding the parametric form of the densities and a transition between continous (differential) and discrete (Shannon) entropies. We can draw strong conceptual parallels between our models and PixelCNN autoencoders , and especially PixelVAE and VLAE , when applied to discrete latents. These models are often evaluated by comparing average likelihoods (which correspond to differential entropies), whereas compression models are typically evaluated by comparing several bit rates (corresponding to Shannon entropies) and distortion values across the rate–distortion frontier, making evaluations more complex.
Discussion
Our approach extends the work of Ballé et al. in two ways. First, we generalize the GSM model to a conditional Gaussian mixture model (GMM). Supporting this model is simply a matter of generating both a mean and a scale parameter conditioned on the hyperprior. Intuitively, the average likelihood of the observed latents increases when the center of the conditional Gaussian is closer to the true value and a smaller scale is predicted, i.e., more structure can be exploited by modeling conditional means. The core question is whether or not the benefits of this more sophisticated model outweigh the cost of the associated side information. We showed in Figure 2 (right) that a GMM-based entropy model provides a net benefit and outperforms the simpler GSM-based model in terms of rate–distortion performance without increasing the asymptotic complexity of the model.
The second extension that we explore is the idea of combining an autoregressive model with the hyperprior. Intuitively, we can see how these components are complementary in two ways. First, starting from the perspective of the hyperprior, we see that for identical hyper-network architectures, improvements to the entropy model require more side information. The side information increases the total compressed file size, which limits its benefit. In contrast, introducing an autoregressive component into the prior does not incur a potential rate penalty since the predictions are based only on the causal context, i.e., on latents that have already been decoded. Similarly, from the perspective of the autoregressive model, we expect some amount of uncertainty that can not be eliminated solely from the causal context. The hyperprior, however, can “look into the future” since it is part of the compressed bitstream and is fully known by the decoder. The hyperprior can thus learn to store information needed to reduce the uncertainty in the autoregressive model while avoiding information that can be accurately predicted from context.
Figure 6 visualizes some of the internal mechanisms of our models. We show three of the variants: one Gaussian scale mixture equivalent to , another strictly hierarchical prior extended to a Gaussian mixture model, and one combined model using an autoregressive component and a hyperprior. After encoding the lighthouse image shown in Figure 5, we extracted the latents for the channel with the highest entropy. These latents are visualized in the first column of Figure 6. The second column holds the conditional means and clearly shows the added detail attained with an autoregressive component, which is reminiscent of the observation that VAE-based models tend to produce blurrier images than autoregressive models . This improvement leads to a lower prediction error (third column) and smaller predicted scales, i.e. smaller uncertainty (fourth column). Our entropy model assumes that latents are conditionally independent given the hyperprior, which implies that the normalized latents, i.e. values with the predicted mean and scale removed, should be closer to i.i.d. Gaussian noise. The fifth column of Figure 6 shows that the combined model is closest to this ideal and that the autoregressive model helps significantly (compare row 4 with row 2). Finally, the last two columns show how the entropy is distributed across the image for the latents and hyper-latents.
From a practical standpoint, autoregressive models are less desirable than hierarchical models since they are inherently serial, and therefore can not be sped up using techniques such as parallelization. To report the performance of the compression models which contain an autoregressive component, we refrained from implementing a full decoder for this paper, and instead compare Shannon entropies. We have empirically verified that these measurements are within a fraction of a percent of the size of the bitstream generated by arithmetic coding.
Probability density distillation has been successfully used to get around the serial nature of autoregressive models for the task of speech synthesis , but unfortunately the same type of method cannot be applied in the domain of compression due to the coupling between the prior and the arithmetic decoder. To address these computational concerns, we have begun to explore very lightweight context models as described in Section 3 and Figure 4, and are considering further techniques to reduce the computational requirements of the Context Model and Entropy Parameters networks, such as engineering a tight integration of the arithmetic decoder with a differentiable autoregressive model. An alternative direction for future research may be to avoid the causality issue altogether by introducing yet more complexity into strictly hierarchical priors.