Enhanced Invertible Encoding for Learned Image Compression

Yueqi Xie, Ka Leong Cheng, Qifeng Chen

Introduction

Lossy image compression has been a fundamental and important research topic in media storage and transmission for decades. Classical image compression standards are usually based on handcrafted schemes, such as JPEG (Wallace, 1992), JPEG2000 (Rabbani, 2002), WebP (Google, 2010), BPG (Bellard, 2015), and Versatile Video Coding (VVC) ((JVET), 2021). Some of them are widely used in practice. Recently, there is an increasing interest in learned image compression methods (Ballé et al., 2016; Minnen et al., 2018; Mentzer et al., 2018; Liu et al., 2019; Cheng et al., 2020) considering their competitive performance. Generally, the recent VAE-based methods (Ballé et al., 2016, 2017) follow a process: an encoder transforms the original pixels x\mathbf{x} into lower-dimensional latent features y\mathbf{y} and then quantizes it to y^\mathbf{\hat{y}} , which can be losslessly compressed using entropy coding methods, like arithmetic coding (Witten et al., 1987; Rissanen and Langdon, 1981). A jointly optimized decoder is utilized to transform y^\mathbf{\hat{y}} back to the reconstructed original image x^\mathbf{\hat{x}}.

Most of the recent developments on this topic focus on the improvement of the entropy model. Ballé et al. (Ballé et al., 2018) propose a variational image compression method with a scale hyperprior. Afterward, Minnen et al. (Minnen et al., 2018) combine the context model (Lee et al., 2019) and an autoregressive model over the latent features. Cheng et al. (Cheng et al., 2020) further improve the method with attention modules and employ discretized Gaussian mixture likelihoods to better parameterize the latent features. These recent improvements on entropy models greatly contribute to efficient compression. However, they usually model the transformation between the image space and the latent feature space using an autoencoder framework. Although autoencoders have strong capacities to select the important portions of the information for reconstruction, the neglected information during encoding is usually lost and unrecoverable for decoding.

To resolve the information loss problem, a possibly good choice is to integrate the idea of invertible neural networks (INNs) into the encoding-decoding process because INNs have the strictly invertible property to help preserve information. However, it is still nontrivial to leverage INNs for replacing the original encoder-decoder architecture. On the one hand, to ensure strict invertibility, the total pixel number in the output should be the same as the input pixel number. It is intractable to quantize and compress such high-dimensional features (Gersho and Gray, 2012) and hard to get desirable low bit rates for compression. The work (Helminger et al., 2021) explores the usage of INN to break AE limit but only gets good performance in high bpp range. In the work (Wang et al., 2020) attempting to use INNs for lower-bpp image compression, they take a subspace of the output to get lower-dimensional features, model lost information with a given distribution, and do sampling during the inverse passing. However, this strategy introduces unwanted errors and leads to unstable training, so they further introduce a knowledge distillation module to guide and stabilize the training of INNs, but it may lead to sub-optimal solutions. On the other hand, although the mathematical network design guarantees the invertibility of INNs, it makes INNs have limited nonlinear transformation capacity (Dinh et al., 2015) compared to other encoder-decoder architectures.

Concerning these aspects, we propose an enhanced Invertible Encoding Network for image compression, which maintains a highly invertible structure based on INN. Instead of using the unstable sampling mechanism for training (Wang et al., 2020), we propose an attentive channel squeeze layer to stabilize the training and to flexibly adjust the feature dimension for a lower bit rate. We also present a feature enhancement module to improve the network nonlinear representation capacity, where the same-resolution transformation and the residual connection of this module help preserve image information. With our design, we can directly integrate INNs to capture the lost information without relying on any sub-optimal guidance and boost performance.

Experimental results show that our proposed method outperforms the existing learned methods and traditional codecs on three widely-used datasets for image compression, including Kodak (Company, 1999) and two high-resolution datasets, namely the CLIC Professional Validation dataset (Toderici et al., 2021) and the Tecnick dataset (Asuni and Giachetti, 2014). Visual results demonstrate that our reconstructed images contain more details under similar BPP rates, beneficial from our highly invertible design. Further analysis of our proposed attentive channel squeeze layer and feature enhancement module verifies their functionality. The contributions of this paper can be summarized as follows:

Unlike widely-used autoencoder style networks for learned image compression, we present an enhanced Invertible Encoding Network with an INN architecture for the transformation between image space and feature space to largely mitigate the information loss problem.

The proposed method outperforms the existing learned image compression methods and traditional compression codecs, including the latest VVC (VTM 12.1), especially on two high-resolution datasets.

We propose an attentive channel squeeze layer to resolve the unstable and sub-optimal training of the INN-based networks for image compression. We further leverage a feature enhancement module to improve the nonlinear representation capacity of our network.

Related Work

Traditional methods. Lossy image compression has long been an important and fundamental topic in image processing. Traditional compression standards includes JPEG (Wallace, 1992), JPEG2000 (Rabbani, 2002), WebP (Google, 2010), Better Portable Graphics (BPG) (Bellard, 2015) and Versatile Video Coding (VVC) ((JVET), 2021). Many of them are widely used in practice. Typically, they follow the pipeline of transformation, quantization, and entropy coding. The transformation is usually based on handcrafted modules with prior knowledge, such as discrete cosine transformation (DCT) (Ahmed et al., 1974) and discrete wavelet transform (DWT) (Marpe et al., 2003). The entropy coders include Huffman coder and some arithmetic coding methods (Rissanen and Langdon, 1981; Witten et al., 1987; Marpe et al., 2003). Some modern standards like BPG and VVC further introduce intra prediction for better compression performance. However, since these traditional methods generally perform compression based on image blocks, their reconstructed images usually contain some limitations of the blocking effects.

Learned methods. In these years, deep learning based methods have raised great interest with impressive performance. These methods try to employ neural networks instead of handcrafted rules for the nonlinear transformation learning between the image space and the latent feature space.

Several recurrent neural network (RNN) based methods (Toderici et al., 2015, 2017; Johnston et al., 2018) progressively encode the residual information from the previous step to compress the image. However, these RNN-based models rely on binary representation at each iteration and cannot directly optimize the rate during training.

Another branch of the methods is based on variational autoencoders (VAEs). Several early works (Theis et al., 2017; Ballé et al., 2017; Agustsson et al., 2017) solve the problem of non-differential quantization and estimation of bit rates, which lay the foundation for end-to-end optimization to minimize the estimated bit rates and the reconstructed image distortion. Afterward, the improvement is mainly in two directions. One direction is to build a more effective entropy model for rate estimation. The work (Ballé et al., 2018) proposes a hyperprior entropy model to use additional bits to capture the distribution of the latent features. Some follow-up methods (Minnen et al., 2018; Lee et al., 2019; Mentzer et al., 2018) integrate context factors to further reduce the spatial redundancy within the latent features. Also, 3D context entropy model (Guo et al., 2020), channel-wise models (Minnen and Singh, 2020) and hierarchical entropy models (Minnen et al., 2018; Hu et al., 2020) are used to better extract correlations of latent features. Cheng et al. (Cheng et al., 2020) propose Gaussian mixture likelihood to replace the original single Gaussian model to improve the accuracy. Another direction is to improve the VAE architecture. CNNs with generalized divisive normalization (GDN) layers (Ballé et al., 2016, 2017) achieve good performance for image compression. Attention mechanism and residual blocks (Liu et al., 2019; Zhou et al., 2019; Zhang et al., 2019; Cheng et al., 2020) are incorporated into the VAE architecture. Some other progress includes generative adversarial training (Rippel and Bourdev, 2017; Santurkar et al., 2018; Agustsson et al., 2019), spatial RNNs (Lin et al., 2020) and multi-scale fusion (Rippel and Bourdev, 2017).

Generally, the VAE-based methods account for a more significant proportion among all the existing learned methods for image compression, considering the great performance and stability of the encoder and decoder architecture. Still, they cannot explicitly solve the information loss problem during encoding, making the neglected information usually unrecoverable for decoding.

2. Invertible Neural Networks

Invertible neural networks (INNs) (Dinh et al., 2015, 2017; Kingma and Dhariwal, 2018) are popular with three great properties of the design: (i) INNs maintain a bijective mapping between inputs and outputs; (ii) the bijective mapping is efficiently accessible; (iii) there exists a tractable Jacobian of the bijective mapping to compute posterior probabilities explicitly. Many recent works in different areas with INN architecture achieve better performance than those with autoencoder style frameworks, especially for the tasks with inherent invertibility.

RealNVP (Dinh et al., 2017) uses a multi-scale architecture with coupling layers and convolution operations to first deal with image processing tasks. Ardizzone et al. (Ardizzone et al., 2019) demonstrate the effectiveness of INNs on both the synthetic data and two real-world applications in the medicine and astrophysics fields. Recently, many works start to use normalizing flow methods with exact likelihoods during training in generative tasks, where the key is to parametrize the distribution using INNs. SRFlow (Lugmayr et al., 2020) uses a conditional INN architecture to better resolve the ill-posed problem of super-resolution compared to GAN-based methods. Pumarola et al. (Pumarola et al., 2020) propose a conditional generative flow model for Image and 3D Point Cloud generation. For the image rescaling task, Xiao et al. (Xiao et al., 2020) use INNs to generate a bijective mapping between high-resolution images and low-resolution images with additional latent variables. Xing et al. (Xing et al., 2021) propose an invertible image signal processing pipeline based on INNs.

Method

Figure 1 provides a high-level overview of general learned image compression models in the transform coding approach (Goyal, 2001). A baseline model (Figure 1(a)) is formulated as follows:

For encoding, a parametric analysis transform gag_{a} encodes the image x\mathbf{x} into latent features y\mathbf{y}, and y\mathbf{y} is then quantized to obtain the discrete latent features y^\mathbf{\hat{y}}, which is then losslessly compressed into bitstreams using entropy coding like arithmetic coding (Rissanen and Langdon, 1981). Note that since integer rounding is a fundamentally non-differentiable function, we use the method in (Ballé et al., 2017) to approximately model the quantized latent features by adding a uniform noise U(−0.5,0.5)U(-0.5,0.5) to y\mathbf{y} during training. For simplicity, we use y^\mathbf{\hat{y}} to denote both the latent features with uniform noise added during training and the discretely quantized latent features during testing. For decoding, a parametric synthesis transform gsg_{s} decodes the quantized y^\mathbf{\hat{y}} back to the reconstructed image x^\mathbf{\hat{x}}.

The fundamental optimization objective of learned image compression is to minimize a weighted sum of the rate-distortion tradeoff during training:

The rate RR is the the entropy of y^\mathbf{\hat{y}}, which is estimated by a non-parametric, fully factorized density model (entropy model) py^∣θp_{\mathbf{\hat{y}}|\theta} during training. So, we have the rate estimation:

The distortion DD is defined as D=MSE(x,x^)D=MSE(\mathbf{x},\mathbf{\hat{x}}) for MSE optimization and D=1−MS-SSIM(x,x^)D=1-MS\text{-}SSIM(\mathbf{x},\mathbf{\hat{x}}) for MS-SSIM (Wang et al., 2003) optimization. We use λ\lambda to control the rate-distortion tradeoff for different bit rates.

Later, Ballé et al. (Ballé et al., 2018) propose a scale hyperprior on top of the general learn image compression model, as shown in Figure 1(b). Specifically, they stack another parametric analysis transform hah_{a} on top of y\mathbf{y} to capture the significant spatial dependencies in the quantized latent features y^\mathbf{\hat{y}} and obtain an additional set of encoded variables z\mathbf{z}. In this way, they model y^\mathbf{\hat{y}} as a zero-mean Gaussian distribution with standard deviations σ\mathbf{\sigma} for all the elements. The standard deviations are estimated by another parametric synthesis transform hsh_{s}, which takes the quantized z^\mathbf{\hat{z}} as input and outputs the estimated standard deviations σ^\mathbf{\hat{\sigma}}. So, we have the conditional probability distribution py^∣z^=N(0,σ2)p_{\mathbf{\hat{y}}|\mathbf{\hat{z}}}=\mathcal{N}(0,\mathbf{\sigma}^{2}). Similarly, an entropy model pz^∣θp_{\mathbf{\hat{z}}|\theta} is applied for entropy estimation of z^\mathbf{\hat{z}}. In this case, the rate estimation contains two terms:

Later works further improve the hyperprior to better parameterize the distributions of the quantized latent features with a more accurate and flexible entropy model. Minnen (Minnen et al., 2018) propose an autoregressive context model with a mean and scale hyperprior, as shown in Figure 1(c). Cheng et al. (Cheng et al., 2020) utilize discretized Gaussian mixture likelihoods with attention enhancement to model the distributions of y^\mathbf{\hat{y}}, as shown in Figure 1(d).

2. Proposed Method

Instead of optimizing the parameterization hah_{a}, hsh_{s} of the latent feature distribution, our proposed method focuses on enhancing the analysis gag_{a} and synthesis gsg_{s} transforms between the image space X\mathcal{X} and the latent feature space Y\mathcal{Y}. Considering the natural invertibility in image compression, we use an invertible network design with a feature enhancement module and an attentive channel squeeze layer to play the role of the analysis gag_{a} and synthesis gsg_{s} transforms. Figure 2 shows an overview of our proposed approach.

INN architecture. We formulate an INN architecture design to serve as the analysis gag_{a} and synthesis gsg_{s} transforms. It consists of two essential invertible layers: the downscaling layer and the coupling layer. Existing methods (Ballé et al., 2018; Cheng et al., 2020; Minnen et al., 2018) usually employ 44 downscaling and 44 corresponding upscaling modules in the analysis and synthesis transforms, respectively. Similarly, we also stack 44 invertible blocks in our INN architecture to downscale the input resolution by a factor of 242^{4}, where each block sequentially contains 11 downscaling layer and 33 coupling layers. The kernel sizes of the coupling layers for the 44 blocks are empirically set as k=5,5,3,3k=5,5,3,3.

The downscaling layer is composed of a pixel shuffling layer (Dinh et al., 2017) and an invertible 1×11\times 1 convolution (Kingma and Dhariwal, 2018), and each downscaling layer reduces the resolution of the input tensor by 22 and quadruples the channel dimension.

We use the affine coupling layer design introduced in RealNVP (Dinh et al., 2017). For the ithi^{th} coupling layer that takes an input u1:C(i)\mathbf{u}^{(i)}_{1:C} with dimensional size of CC, it splits the input at position c<Cc<C into two parts and gives a CC dimensional output u1:C(i+1)\mathbf{u}^{(i+1)}_{1:C}:

where ⊙\odot denotes the Hadamard product, exp⁡(⋅)\exp(\cdot) denotes the exponential function, and σc(⋅)\sigma_{c}(\cdot) denotes the center sigmoid function. Symmetrically, the ithi^{th} coupling layer inversely takes u1:C(i+1)\mathbf{u}^{(i+1)}_{1:C} as input with splitting position cc. The affine coupling layer gives a perfect inverse:

Please note that such invertibility is inherently guaranteed by the mathematical design. Thus, the invertibility holds for any arbitrary feedforward functions g1,g2,h1,h2g_{1},g_{2},h_{1},h_{2}, meaning that these functions need not be invertible. In our implementations, all these four functions use the following bottleneck design: Conv-LeakyReLU-Conv-LeakyReLU-Conv, where the kernel size of the first and the last convolution is set as kk; the middle one is a convolutional layer with filter size 11.

Feature enhancement module. INNs are powerful when modeling invertible transformation. However, since the property of invertibility is guaranteed by the strictly invertible network design, INNs usually have a limited capacity of nonlinear representation (Dinh et al., 2015). Hence, we add a feature enhancement module in a residual manner before the INN architecture to improve the nonlinear representativeness of our network. Specifically, this module is based on the popular Dense Block (Huang et al., 2017), and “Convs (1,3,1)” means three cascade convolutions with kernel size 11, 33, 11.

Attentive channel squeeze. All the operations in INNs cannot change the total number of pixels in the input tensor, many of which are simply redundant pixels to be compressed. To resolve this problem, we introduce an attentive channel squeeze layer to reduce the channel dimension of the output tensor of INNs. The idea is as follows: Given a compression ratio α\alpha and an input tensor v\mathbf{v} with size (CC, HH, WW), the attentive channel squeeze layer first forwardly reshape the tensor into a shape of (α\alpha, Cα\frac{C}{\alpha}, HH, WW). Then it performs average operation along the first dimension to obtain the latent features y\mathbf{y} with size (Cα\frac{C}{\alpha}, HH, WW), followed by an attention module proposed in (Cheng et al., 2020). For the inverse process, the attentive channel squeeze layer makes α\alpha copies of the quantized y^\mathbf{\hat{y}} first after the attention module and reshapes it into a size of (CC, HH, WW).

3. Overall Workflow

For the decoding process, the attentive channel squeeze layer copies the quantized latent features y^\mathbf{\hat{y}} after attention enhancement for α\alpha times and then reshapes it as v^\mathbf{\hat{v}}. The inverse pass of the INN architecture decodes v^\mathbf{\hat{v}} to obtain u^\mathbf{\hat{u}}, which finally undergoes the feature enhancement module for the reconstructed image x^\mathbf{\hat{x}}.

We employ the same hyperprior proposed by Minnen et al. (Minnen et al., 2018), which uses a mean and scale Gaussian distribution to parameterize the quantized latent features y^\mathbf{\hat{y}} with a pair of analysis hah_{a} and synthesis hsh_{s} transforms. Specifically, hah_{a} obtains the side information z=ha(y^)\mathbf{z}=h_{a}(\mathbf{\hat{y}}), and hsh_{s} takes the quantized side information z^\mathbf{\hat{z}} as input. The output hs(z^)h_{s}(\mathbf{\hat{z}}) works with the causal context Cm(y^)C_{m}(\mathbf{\hat{y}}) for the mean and scale Gaussian model py^∣z^←hs(z^),Cm(y^)p_{\mathbf{\hat{y}}|\mathbf{\hat{z}}}\leftarrow h_{s}(\mathbf{\hat{z}}),C_{m}(\mathbf{\hat{y}}). Since there is no prior for z^\mathbf{\hat{z}}, a factorized-prior entropy model θ\theta introduced in (Ballé et al., 2017) is used to parametize the distribution of z^\mathbf{\hat{z}} as pz^∣θp_{\mathbf{\hat{z}}|\theta}. For entropy coding, we use the asymmetric numeral systems (ANS) (Duda, 2009) to losslessly compress y^\mathbf{\hat{y}} and z^\mathbf{\hat{z}} into bitstreams.

Experiments

We conduct experiments on three commonly used image compression datasets to validate our method. In the remainder of this section, we first introduce the detailed experimental setup, followed by quantitative and qualitative comparisons with existing state-of-the-art methods. We further conduct some analysis and ablation studies to examine the effectiveness of our proposed feature enhancement module and attentive channel squeeze layer.

Training details. We use the Flicker 2W dataset used in (Liu et al., 2020), consisting of 20,74520,745 high-quality general images for the image compression task. We randomly select around 200200 images for our validation set, and the remaining images are used for training. Our network is trained on 256×256256\times 256 randomly cropped patches, using the recently well-developed CompressAI PyTorch library (Bégaint et al., 2020). Note that we drop a few images with either height or width smaller than 256256 px for convenience.

All the experiments are conducted on a single RTX 2080 Ti GPU and trained for 600600 epochs with a batch size of 88 using Adam (Kingma and Ba, 2015) optimizer. Generally, it takes around 1010 days to train a model. Our network is first optimized for 450450 epochs with an initial learning rate of 10−410^{-4}; then the learning rate is reduced to 10−510^{-5} at epoch 450450 and further down to 10−610^{-6} at epoch 550550.

We use the channel number N=CαN=\frac{C}{\alpha} and the weight factor λ\lambda as our quality parameters. We in-total train 88 models optimized with MSE (mean squared error) quality metric and 55 models with MS-SSIM (multiscale structural similarity) quality metric (Wang et al., 2003) under different quality levels. For the MSE models, λ\lambda is chosen from the set {0.0016,0.0024,0.0032,0.0075,0.015,0.03,0.045,0.09}\{0.0016,0.0024,0.0032,0.0075,0.015,0.03,0.045,0.09\}, in which the first four λ\lambda values are paired with channel number N=128N=128 for lower-rate models, and NN is set as 192192 to pair with the remaining four λ\lambda values for higher-rate models. For the MS-SSIM models, λ\lambda belongs to the set of {6,12,40,120,220}\{6,12,40,120,220\}, and we use N=128N=128 for the first two λ\lambda values to train lower-rate models and N=192N=192 for the remaining three λ\lambda values to train higher-rate models.

Evaluation. We evaluate our methods on three commonly used datasets for image compression, which are the Kodak PhotoCD image dataset (Kodak) (Company, 1999), the CLIC Professional Validation dataset (CLIC) (Toderici et al., 2021), and the old Tecnick dataset (Asuni and Giachetti, 2014). The Kodak dataset contains 2424 uncompressed images with resolutions of 768×512768\times 512; the CLIC dataset comprises 4141 high-quality images with much higher resolutions; the Tecnick dataset contains 100100 images with high resolutions of 1200×12001200\times 1200.

We use the peak signal-to-noise ratio (PSNR) and the multiscale structural similarity index (MS-SSIM) (Wang et al., 2003) to quantify the image distortion level; we use the bits per pixel (bpp) to evaluate the rate performance. We draw the rate-distortion (RD) curves according to their rate-distortion performance to compare the coding efficiency of different methods. We also report the area under the rate-distortion curve (AUC) as an aggregate measurement to better compare methods with similar performance.

2. Rate-distortion Performance

We compare our model with state-of-the-art learned image compression models, including the methods proposed by Ballé el al. (Ballé et al., 2018), Minnen et al. (Minnen et al., 2018), Lee et al. (Lee et al., 2019), Hu et al. (Hu et al., 2020), and Cheng et al. (Cheng et al., 2020). The corresponding data points on their RD curves are collected from their paper or their official GitHub pages. We also compare with some widely-used image compression codecs, including JPEG (Wallace, 1992), JPEG2000 (Rabbani, 2002), WebP (Google, 2010), BPG (Bellard, 2015), and VVC ((JVET), 2021). We evaluate their performance using the CompressAI evaluation platform. For VVC, we use the most up-to-date VVC Official Test Model VTM 12.1 (accessed on April 2021) with an intra-profile configuration from the official GitHub page to test on images. To evaluate the BPG performance, we use the BPG software with subsampling mode of YUV444, HEVC implementation of x265, and bit depth of 88 to test on images.

Figure 3 shows the RD curve comparison on the Kodak dataset. Similar to existing work (Cheng et al., 2020), we convert MS−SSIMMS-SSIM to −10log⁡10(1−MS-SSIM)-10\log_{10}(1-MS\text{-}SSIM) for clearer comparison. It can be observed that our method slightly outperforms VVC (VTM 12.1) and yields much better performance when comparing with both the existing learned methods and other traditional image compression standards.

For the CLIC dataset and the Tecnick dataset, we compare our MSE optimized results with traditional compression standards and the learned methods with official testing results available in their paper or their official GitHub pages. We show the RD curves on the CLIC dataset in Figure 4 and the Tecnick dataset in Figure 5. We can see that our MSE optimized method outperforms all other approaches. Note that most images in the CLIC dataset and the Tecknick dataset are of high resolutions, implying that our method is more robust and promising to compress high-resolution images.

From the RD curves, we can see that VVC (VTM 12.1) and our approach achieve similar performance in terms of PSNR. Hence, we further report their corresponding AUC values for better comparison and ranking. Statistics in Table 1 indicates that our method outperforms the latest VVC (VTM 12.1) codec in terms of the aggregated AUC metric.

3. Qualitative Results

We show some qualitative comparison of some sample reconstructed images on the Kodak dataset in Figure 6. The first sample is image kodim01 with an approximate bpp of 0.1450.145; the second sample is image kodim07 with approximately 0.1250.125 bpp; the last sample is image kodim22 with around 0.1300.130 bpp. For JPEG and JPEG2000, we use the lowest quality since they cannot reach the mentioned bpp levels. We can see that our MSE optimized method achieves a good performance compared with the latest VVC (VTM 12.1) codec, much better than the performance of other codecs. Also, our MS-SSIM optimized method can reconstruct images with much more structural details than all the traditional codecs. We present more qualitative results in our supplementary materials.

4. Analysis of Attentive Channel Squeeze

In this section, we demonstrate that our proposed attentive channel squeeze layer introduces only minor deviation to enable a stable and tractable dimension adjustment. We also analyze the distribution of the deviation maps on images and the deviation levels of different quality models.

To quantify such deviation, we do some analysis using the Kodak dataset to calculate the mean absolute pixel deviation as follows:

where γi,j\gamma_{i,j} and γ^1,j\hat{\gamma}_{1,j} denote the corresponding pixel in γ\mathbf{\gamma} and γ^\mathbf{\hat{\gamma}}, respectively. As shown in Table 2, the value of the deviation is minor, which is less than or comparable with the error due to the quantization in most cases.

To compare the deviation among models under different quality levels, it is unfair to directly compare the values because γ\mathbf{\gamma} and γ^\mathbf{\hat{\gamma}} have a larger value range for the model with higher quality. Accordingly, the absolute value of deviation is amplified. As a result, calculating the “relative” deviation concerning the value range is a more reasonable practice. A naive way is to divide the mean absolute pixel deviation ϵ\epsilon by the value range for fair comparison among different quality models. However, we find that the distribution of the pixel values has long tails, so the value range is very likely to be affected by the outlier pixel values. Hence, to better scale the mean absolute pixel deviation, we introduce a scaling factor μ\mu:

We further visualize the scaled deviation map of the image kodim20 and image kodim24 between γ\mathbf{\gamma} and γ^\mathbf{\hat{\gamma}} of our models under different quality levels in Figure 7. Note that the deviation map is in shape (h,w)(h,w), and each pixel is the mean of the absolute deviation along the channel dimension after scaling with μ\mu. We can see that the “relative” deviation generally decreases as the quality increases, indicating less information loss of the averaging operation, which is in line with our intuition: higher-quality models usually lead to less information loss in the process of compression.

5. Ablation Study

To verify the contribution of the proposed nonlinear feature enhancement module, we conduct a corresponding ablation study on this module. Specifically, we train two models with and without using the nonlinear feature enhancement module on the Flicker 2W dataset for 600600 epochs, using the same weight factor λ=0.01\lambda=0.01 and channel number N=192N=192.

Figure 8 shows the rate-distortion points of two models evaluated on the Kodak dataset. We can observe that the proposed nonlinear feature enhancement module improves the modeling capacity of the model to compress images with lower bit rates and higher PSNR and MS-SSIM values.

Conclusion

Unlike existing autoencoder style networks, our proposed enhanced Invertible Encoding Network based on invertible neural networks (INNs) can better model image compression as an invertible process. The critical issues in integrating INNs for image compression include unstable training and limited nonlinear transformation capacity. Our proposed attentive channel squeeze layer offers stable and tractable feature dimension adjustment; the incorporated feature enhancement module increases the network nonlinear transformation capacity. Overall, our network maintains a highly invertible architecture to largely mitigate information loss when compressing images.

Extensive experiments on three widely-used datasets show that our approach outperforms state-of-the-art learned image compression methods and existing compression standards, including VVC (VTM 12.1), especially on two high-resolution image datasets. Furthermore, the visual results demonstrate that our compression methods preserve more detailed information than current compression standards.

References