Enhanced Invertible Encoding for Learned Image Compression
Yueqi Xie, Ka Leong Cheng, Qifeng Chen
Introduction
Lossy image compression has been a fundamental and important research topic in media storage and transmission for decades. Classical image compression standards are usually based on handcrafted schemes, such as JPEG (Wallace, 1992), JPEG2000 (Rabbani, 2002), WebP (Google, 2010), BPG (Bellard, 2015), and Versatile Video Coding (VVC) ((JVET), 2021). Some of them are widely used in practice. Recently, there is an increasing interest in learned image compression methods (Ballé et al., 2016; Minnen et al., 2018; Mentzer et al., 2018; Liu et al., 2019; Cheng et al., 2020) considering their competitive performance. Generally, the recent VAE-based methods (Ballé et al., 2016, 2017) follow a process: an encoder transforms the original pixels into lower-dimensional latent features and then quantizes it to , which can be losslessly compressed using entropy coding methods, like arithmetic coding (Witten et al., 1987; Rissanen and Langdon, 1981). A jointly optimized decoder is utilized to transform back to the reconstructed original image .
Most of the recent developments on this topic focus on the improvement of the entropy model. Ballé et al. (Ballé et al., 2018) propose a variational image compression method with a scale hyperprior. Afterward, Minnen et al. (Minnen et al., 2018) combine the context model (Lee et al., 2019) and an autoregressive model over the latent features. Cheng et al. (Cheng et al., 2020) further improve the method with attention modules and employ discretized Gaussian mixture likelihoods to better parameterize the latent features. These recent improvements on entropy models greatly contribute to efficient compression. However, they usually model the transformation between the image space and the latent feature space using an autoencoder framework. Although autoencoders have strong capacities to select the important portions of the information for reconstruction, the neglected information during encoding is usually lost and unrecoverable for decoding.
To resolve the information loss problem, a possibly good choice is to integrate the idea of invertible neural networks (INNs) into the encoding-decoding process because INNs have the strictly invertible property to help preserve information. However, it is still nontrivial to leverage INNs for replacing the original encoder-decoder architecture. On the one hand, to ensure strict invertibility, the total pixel number in the output should be the same as the input pixel number. It is intractable to quantize and compress such high-dimensional features (Gersho and Gray, 2012) and hard to get desirable low bit rates for compression. The work (Helminger et al., 2021) explores the usage of INN to break AE limit but only gets good performance in high bpp range. In the work (Wang et al., 2020) attempting to use INNs for lower-bpp image compression, they take a subspace of the output to get lower-dimensional features, model lost information with a given distribution, and do sampling during the inverse passing. However, this strategy introduces unwanted errors and leads to unstable training, so they further introduce a knowledge distillation module to guide and stabilize the training of INNs, but it may lead to sub-optimal solutions. On the other hand, although the mathematical network design guarantees the invertibility of INNs, it makes INNs have limited nonlinear transformation capacity (Dinh et al., 2015) compared to other encoder-decoder architectures.
Concerning these aspects, we propose an enhanced Invertible Encoding Network for image compression, which maintains a highly invertible structure based on INN. Instead of using the unstable sampling mechanism for training (Wang et al., 2020), we propose an attentive channel squeeze layer to stabilize the training and to flexibly adjust the feature dimension for a lower bit rate. We also present a feature enhancement module to improve the network nonlinear representation capacity, where the same-resolution transformation and the residual connection of this module help preserve image information. With our design, we can directly integrate INNs to capture the lost information without relying on any sub-optimal guidance and boost performance.
Experimental results show that our proposed method outperforms the existing learned methods and traditional codecs on three widely-used datasets for image compression, including Kodak (Company, 1999) and two high-resolution datasets, namely the CLIC Professional Validation dataset (Toderici et al., 2021) and the Tecnick dataset (Asuni and Giachetti, 2014). Visual results demonstrate that our reconstructed images contain more details under similar BPP rates, beneficial from our highly invertible design. Further analysis of our proposed attentive channel squeeze layer and feature enhancement module verifies their functionality. The contributions of this paper can be summarized as follows:
Unlike widely-used autoencoder style networks for learned image compression, we present an enhanced Invertible Encoding Network with an INN architecture for the transformation between image space and feature space to largely mitigate the information loss problem.
The proposed method outperforms the existing learned image compression methods and traditional compression codecs, including the latest VVC (VTM 12.1), especially on two high-resolution datasets.
We propose an attentive channel squeeze layer to resolve the unstable and sub-optimal training of the INN-based networks for image compression. We further leverage a feature enhancement module to improve the nonlinear representation capacity of our network.
Related Work
Traditional methods. Lossy image compression has long been an important and fundamental topic in image processing. Traditional compression standards includes JPEG (Wallace, 1992), JPEG2000 (Rabbani, 2002), WebP (Google, 2010), Better Portable Graphics (BPG) (Bellard, 2015) and Versatile Video Coding (VVC) ((JVET), 2021). Many of them are widely used in practice. Typically, they follow the pipeline of transformation, quantization, and entropy coding. The transformation is usually based on handcrafted modules with prior knowledge, such as discrete cosine transformation (DCT) (Ahmed et al., 1974) and discrete wavelet transform (DWT) (Marpe et al., 2003). The entropy coders include Huffman coder and some arithmetic coding methods (Rissanen and Langdon, 1981; Witten et al., 1987; Marpe et al., 2003). Some modern standards like BPG and VVC further introduce intra prediction for better compression performance. However, since these traditional methods generally perform compression based on image blocks, their reconstructed images usually contain some limitations of the blocking effects.
Learned methods. In these years, deep learning based methods have raised great interest with impressive performance. These methods try to employ neural networks instead of handcrafted rules for the nonlinear transformation learning between the image space and the latent feature space.
Several recurrent neural network (RNN) based methods (Toderici et al., 2015, 2017; Johnston et al., 2018) progressively encode the residual information from the previous step to compress the image. However, these RNN-based models rely on binary representation at each iteration and cannot directly optimize the rate during training.
Another branch of the methods is based on variational autoencoders (VAEs). Several early works (Theis et al., 2017; Ballé et al., 2017; Agustsson et al., 2017) solve the problem of non-differential quantization and estimation of bit rates, which lay the foundation for end-to-end optimization to minimize the estimated bit rates and the reconstructed image distortion. Afterward, the improvement is mainly in two directions. One direction is to build a more effective entropy model for rate estimation. The work (Ballé et al., 2018) proposes a hyperprior entropy model to use additional bits to capture the distribution of the latent features. Some follow-up methods (Minnen et al., 2018; Lee et al., 2019; Mentzer et al., 2018) integrate context factors to further reduce the spatial redundancy within the latent features. Also, 3D context entropy model (Guo et al., 2020), channel-wise models (Minnen and Singh, 2020) and hierarchical entropy models (Minnen et al., 2018; Hu et al., 2020) are used to better extract correlations of latent features. Cheng et al. (Cheng et al., 2020) propose Gaussian mixture likelihood to replace the original single Gaussian model to improve the accuracy. Another direction is to improve the VAE architecture. CNNs with generalized divisive normalization (GDN) layers (Ballé et al., 2016, 2017) achieve good performance for image compression. Attention mechanism and residual blocks (Liu et al., 2019; Zhou et al., 2019; Zhang et al., 2019; Cheng et al., 2020) are incorporated into the VAE architecture. Some other progress includes generative adversarial training (Rippel and Bourdev, 2017; Santurkar et al., 2018; Agustsson et al., 2019), spatial RNNs (Lin et al., 2020) and multi-scale fusion (Rippel and Bourdev, 2017).
Generally, the VAE-based methods account for a more significant proportion among all the existing learned methods for image compression, considering the great performance and stability of the encoder and decoder architecture. Still, they cannot explicitly solve the information loss problem during encoding, making the neglected information usually unrecoverable for decoding.
2. Invertible Neural Networks
Invertible neural networks (INNs) (Dinh et al., 2015, 2017; Kingma and Dhariwal, 2018) are popular with three great properties of the design: (i) INNs maintain a bijective mapping between inputs and outputs; (ii) the bijective mapping is efficiently accessible; (iii) there exists a tractable Jacobian of the bijective mapping to compute posterior probabilities explicitly. Many recent works in different areas with INN architecture achieve better performance than those with autoencoder style frameworks, especially for the tasks with inherent invertibility.
RealNVP (Dinh et al., 2017) uses a multi-scale architecture with coupling layers and convolution operations to first deal with image processing tasks. Ardizzone et al. (Ardizzone et al., 2019) demonstrate the effectiveness of INNs on both the synthetic data and two real-world applications in the medicine and astrophysics fields. Recently, many works start to use normalizing flow methods with exact likelihoods during training in generative tasks, where the key is to parametrize the distribution using INNs. SRFlow (Lugmayr et al., 2020) uses a conditional INN architecture to better resolve the ill-posed problem of super-resolution compared to GAN-based methods. Pumarola et al. (Pumarola et al., 2020) propose a conditional generative flow model for Image and 3D Point Cloud generation. For the image rescaling task, Xiao et al. (Xiao et al., 2020) use INNs to generate a bijective mapping between high-resolution images and low-resolution images with additional latent variables. Xing et al. (Xing et al., 2021) propose an invertible image signal processing pipeline based on INNs.
Method
Figure 1 provides a high-level overview of general learned image compression models in the transform coding approach (Goyal, 2001). A baseline model (Figure 1(a)) is formulated as follows:
For encoding, a parametric analysis transform encodes the image into latent features , and is then quantized to obtain the discrete latent features , which is then losslessly compressed into bitstreams using entropy coding like arithmetic coding (Rissanen and Langdon, 1981). Note that since integer rounding is a fundamentally non-differentiable function, we use the method in (Ballé et al., 2017) to approximately model the quantized latent features by adding a uniform noise to during training. For simplicity, we use to denote both the latent features with uniform noise added during training and the discretely quantized latent features during testing. For decoding, a parametric synthesis transform decodes the quantized back to the reconstructed image .
The fundamental optimization objective of learned image compression is to minimize a weighted sum of the rate-distortion tradeoff during training:
The rate is the the entropy of , which is estimated by a non-parametric, fully factorized density model (entropy model) during training. So, we have the rate estimation:
The distortion is defined as for MSE optimization and for MS-SSIM (Wang et al., 2003) optimization. We use to control the rate-distortion tradeoff for different bit rates.
Later, Ballé et al. (Ballé et al., 2018) propose a scale hyperprior on top of the general learn image compression model, as shown in Figure 1(b). Specifically, they stack another parametric analysis transform on top of to capture the significant spatial dependencies in the quantized latent features and obtain an additional set of encoded variables . In this way, they model as a zero-mean Gaussian distribution with standard deviations for all the elements. The standard deviations are estimated by another parametric synthesis transform , which takes the quantized as input and outputs the estimated standard deviations . So, we have the conditional probability distribution . Similarly, an entropy model is applied for entropy estimation of . In this case, the rate estimation contains two terms:
Later works further improve the hyperprior to better parameterize the distributions of the quantized latent features with a more accurate and flexible entropy model. Minnen (Minnen et al., 2018) propose an autoregressive context model with a mean and scale hyperprior, as shown in Figure 1(c). Cheng et al. (Cheng et al., 2020) utilize discretized Gaussian mixture likelihoods with attention enhancement to model the distributions of , as shown in Figure 1(d).
2. Proposed Method
Instead of optimizing the parameterization , of the latent feature distribution, our proposed method focuses on enhancing the analysis and synthesis transforms between the image space and the latent feature space . Considering the natural invertibility in image compression, we use an invertible network design with a feature enhancement module and an attentive channel squeeze layer to play the role of the analysis and synthesis transforms. Figure 2 shows an overview of our proposed approach.
INN architecture. We formulate an INN architecture design to serve as the analysis and synthesis transforms. It consists of two essential invertible layers: the downscaling layer and the coupling layer. Existing methods (Ballé et al., 2018; Cheng et al., 2020; Minnen et al., 2018) usually employ downscaling and corresponding upscaling modules in the analysis and synthesis transforms, respectively. Similarly, we also stack invertible blocks in our INN architecture to downscale the input resolution by a factor of , where each block sequentially contains downscaling layer and coupling layers. The kernel sizes of the coupling layers for the blocks are empirically set as .
The downscaling layer is composed of a pixel shuffling layer (Dinh et al., 2017) and an invertible convolution (Kingma and Dhariwal, 2018), and each downscaling layer reduces the resolution of the input tensor by and quadruples the channel dimension.
We use the affine coupling layer design introduced in RealNVP (Dinh et al., 2017). For the coupling layer that takes an input with dimensional size of , it splits the input at position into two parts and gives a dimensional output :
where denotes the Hadamard product, denotes the exponential function, and denotes the center sigmoid function. Symmetrically, the coupling layer inversely takes as input with splitting position . The affine coupling layer gives a perfect inverse:
Please note that such invertibility is inherently guaranteed by the mathematical design. Thus, the invertibility holds for any arbitrary feedforward functions , meaning that these functions need not be invertible. In our implementations, all these four functions use the following bottleneck design: Conv-LeakyReLU-Conv-LeakyReLU-Conv, where the kernel size of the first and the last convolution is set as ; the middle one is a convolutional layer with filter size .
Feature enhancement module. INNs are powerful when modeling invertible transformation. However, since the property of invertibility is guaranteed by the strictly invertible network design, INNs usually have a limited capacity of nonlinear representation (Dinh et al., 2015). Hence, we add a feature enhancement module in a residual manner before the INN architecture to improve the nonlinear representativeness of our network. Specifically, this module is based on the popular Dense Block (Huang et al., 2017), and “Convs (1,3,1)” means three cascade convolutions with kernel size , , .
Attentive channel squeeze. All the operations in INNs cannot change the total number of pixels in the input tensor, many of which are simply redundant pixels to be compressed. To resolve this problem, we introduce an attentive channel squeeze layer to reduce the channel dimension of the output tensor of INNs. The idea is as follows: Given a compression ratio and an input tensor with size (, , ), the attentive channel squeeze layer first forwardly reshape the tensor into a shape of (, , , ). Then it performs average operation along the first dimension to obtain the latent features with size (, , ), followed by an attention module proposed in (Cheng et al., 2020). For the inverse process, the attentive channel squeeze layer makes copies of the quantized first after the attention module and reshapes it into a size of (, , ).
3. Overall Workflow
For the decoding process, the attentive channel squeeze layer copies the quantized latent features after attention enhancement for times and then reshapes it as . The inverse pass of the INN architecture decodes to obtain , which finally undergoes the feature enhancement module for the reconstructed image .
We employ the same hyperprior proposed by Minnen et al. (Minnen et al., 2018), which uses a mean and scale Gaussian distribution to parameterize the quantized latent features with a pair of analysis and synthesis transforms. Specifically, obtains the side information , and takes the quantized side information as input. The output works with the causal context for the mean and scale Gaussian model . Since there is no prior for , a factorized-prior entropy model introduced in (Ballé et al., 2017) is used to parametize the distribution of as . For entropy coding, we use the asymmetric numeral systems (ANS) (Duda, 2009) to losslessly compress and into bitstreams.
Experiments
We conduct experiments on three commonly used image compression datasets to validate our method. In the remainder of this section, we first introduce the detailed experimental setup, followed by quantitative and qualitative comparisons with existing state-of-the-art methods. We further conduct some analysis and ablation studies to examine the effectiveness of our proposed feature enhancement module and attentive channel squeeze layer.
Training details. We use the Flicker 2W dataset used in (Liu et al., 2020), consisting of high-quality general images for the image compression task. We randomly select around images for our validation set, and the remaining images are used for training. Our network is trained on randomly cropped patches, using the recently well-developed CompressAI PyTorch library (Bégaint et al., 2020). Note that we drop a few images with either height or width smaller than px for convenience.
All the experiments are conducted on a single RTX 2080 Ti GPU and trained for epochs with a batch size of using Adam (Kingma and Ba, 2015) optimizer. Generally, it takes around days to train a model. Our network is first optimized for epochs with an initial learning rate of ; then the learning rate is reduced to at epoch and further down to at epoch .
We use the channel number and the weight factor as our quality parameters. We in-total train models optimized with MSE (mean squared error) quality metric and models with MS-SSIM (multiscale structural similarity) quality metric (Wang et al., 2003) under different quality levels. For the MSE models, is chosen from the set , in which the first four values are paired with channel number for lower-rate models, and is set as to pair with the remaining four values for higher-rate models. For the MS-SSIM models, belongs to the set of , and we use for the first two values to train lower-rate models and for the remaining three values to train higher-rate models.
Evaluation. We evaluate our methods on three commonly used datasets for image compression, which are the Kodak PhotoCD image dataset (Kodak) (Company, 1999), the CLIC Professional Validation dataset (CLIC) (Toderici et al., 2021), and the old Tecnick dataset (Asuni and Giachetti, 2014). The Kodak dataset contains uncompressed images with resolutions of ; the CLIC dataset comprises high-quality images with much higher resolutions; the Tecnick dataset contains images with high resolutions of .
We use the peak signal-to-noise ratio (PSNR) and the multiscale structural similarity index (MS-SSIM) (Wang et al., 2003) to quantify the image distortion level; we use the bits per pixel (bpp) to evaluate the rate performance. We draw the rate-distortion (RD) curves according to their rate-distortion performance to compare the coding efficiency of different methods. We also report the area under the rate-distortion curve (AUC) as an aggregate measurement to better compare methods with similar performance.
2. Rate-distortion Performance
We compare our model with state-of-the-art learned image compression models, including the methods proposed by Ballé el al. (Ballé et al., 2018), Minnen et al. (Minnen et al., 2018), Lee et al. (Lee et al., 2019), Hu et al. (Hu et al., 2020), and Cheng et al. (Cheng et al., 2020). The corresponding data points on their RD curves are collected from their paper or their official GitHub pages. We also compare with some widely-used image compression codecs, including JPEG (Wallace, 1992), JPEG2000 (Rabbani, 2002), WebP (Google, 2010), BPG (Bellard, 2015), and VVC ((JVET), 2021). We evaluate their performance using the CompressAI evaluation platform. For VVC, we use the most up-to-date VVC Official Test Model VTM 12.1 (accessed on April 2021) with an intra-profile configuration from the official GitHub page to test on images. To evaluate the BPG performance, we use the BPG software with subsampling mode of YUV444, HEVC implementation of x265, and bit depth of to test on images.
Figure 3 shows the RD curve comparison on the Kodak dataset. Similar to existing work (Cheng et al., 2020), we convert to for clearer comparison. It can be observed that our method slightly outperforms VVC (VTM 12.1) and yields much better performance when comparing with both the existing learned methods and other traditional image compression standards.
For the CLIC dataset and the Tecnick dataset, we compare our MSE optimized results with traditional compression standards and the learned methods with official testing results available in their paper or their official GitHub pages. We show the RD curves on the CLIC dataset in Figure 4 and the Tecnick dataset in Figure 5. We can see that our MSE optimized method outperforms all other approaches. Note that most images in the CLIC dataset and the Tecknick dataset are of high resolutions, implying that our method is more robust and promising to compress high-resolution images.
From the RD curves, we can see that VVC (VTM 12.1) and our approach achieve similar performance in terms of PSNR. Hence, we further report their corresponding AUC values for better comparison and ranking. Statistics in Table 1 indicates that our method outperforms the latest VVC (VTM 12.1) codec in terms of the aggregated AUC metric.
3. Qualitative Results
We show some qualitative comparison of some sample reconstructed images on the Kodak dataset in Figure 6. The first sample is image kodim01 with an approximate bpp of ; the second sample is image kodim07 with approximately bpp; the last sample is image kodim22 with around bpp. For JPEG and JPEG2000, we use the lowest quality since they cannot reach the mentioned bpp levels. We can see that our MSE optimized method achieves a good performance compared with the latest VVC (VTM 12.1) codec, much better than the performance of other codecs. Also, our MS-SSIM optimized method can reconstruct images with much more structural details than all the traditional codecs. We present more qualitative results in our supplementary materials.
4. Analysis of Attentive Channel Squeeze
In this section, we demonstrate that our proposed attentive channel squeeze layer introduces only minor deviation to enable a stable and tractable dimension adjustment. We also analyze the distribution of the deviation maps on images and the deviation levels of different quality models.
To quantify such deviation, we do some analysis using the Kodak dataset to calculate the mean absolute pixel deviation as follows:
where and denote the corresponding pixel in and , respectively. As shown in Table 2, the value of the deviation is minor, which is less than or comparable with the error due to the quantization in most cases.
To compare the deviation among models under different quality levels, it is unfair to directly compare the values because and have a larger value range for the model with higher quality. Accordingly, the absolute value of deviation is amplified. As a result, calculating the “relative” deviation concerning the value range is a more reasonable practice. A naive way is to divide the mean absolute pixel deviation by the value range for fair comparison among different quality models. However, we find that the distribution of the pixel values has long tails, so the value range is very likely to be affected by the outlier pixel values. Hence, to better scale the mean absolute pixel deviation, we introduce a scaling factor :
We further visualize the scaled deviation map of the image kodim20 and image kodim24 between and of our models under different quality levels in Figure 7. Note that the deviation map is in shape , and each pixel is the mean of the absolute deviation along the channel dimension after scaling with . We can see that the “relative” deviation generally decreases as the quality increases, indicating less information loss of the averaging operation, which is in line with our intuition: higher-quality models usually lead to less information loss in the process of compression.
5. Ablation Study
To verify the contribution of the proposed nonlinear feature enhancement module, we conduct a corresponding ablation study on this module. Specifically, we train two models with and without using the nonlinear feature enhancement module on the Flicker 2W dataset for epochs, using the same weight factor and channel number .
Figure 8 shows the rate-distortion points of two models evaluated on the Kodak dataset. We can observe that the proposed nonlinear feature enhancement module improves the modeling capacity of the model to compress images with lower bit rates and higher PSNR and MS-SSIM values.
Conclusion
Unlike existing autoencoder style networks, our proposed enhanced Invertible Encoding Network based on invertible neural networks (INNs) can better model image compression as an invertible process. The critical issues in integrating INNs for image compression include unstable training and limited nonlinear transformation capacity. Our proposed attentive channel squeeze layer offers stable and tractable feature dimension adjustment; the incorporated feature enhancement module increases the network nonlinear transformation capacity. Overall, our network maintains a highly invertible architecture to largely mitigate information loss when compressing images.
Extensive experiments on three widely-used datasets show that our approach outperforms state-of-the-art learned image compression methods and existing compression standards, including VVC (VTM 12.1), especially on two high-resolution image datasets. Furthermore, the visual results demonstrate that our compression methods preserve more detailed information than current compression standards.