FVC: A New Framework towards Deep Video Compression in Feature Space

Zhihao Hu, Guo Lu, Dong Xu

Introduction

There is an increasing research interest in developing the next generation video compression technologies. While the traditional video compression methods have achieved promising performance by using the hand-designed modules (e.g., block-based motion estimation, Discrete Cosine Transform (DCT)) to reduce spatial and temporal redundancy, these modules cannot be optimized in an end-to-end fashion based on large-scale video data.

The recent deep video compression works have achieved impressive results by applying the deep neural networks within the hybrid video compression framework. Currently, most works only rely on pixel-level operations (e.g., motion estimation or motion compensation) for reducing redundancy. For example, pixel-level optical flow estimation and motion compensation are used in to reduce temporal redundancy, and pixel-level residual is further compressed by using the auto-encoder style network. However, this pixel-level paradigm suffers from the following three limitations. First, it is difficult to produce accurate pixel-level optical flow information, especially for videos with complex non-rigid motion patterns. Second, even we can extract sufficiently accurate motion information, the pixel-level motion compensation process may introduce additional artifacts. Third, it is also a non-trivial task to compress pixel-level residual information.

Given robust representation ability of deep features for various applications, it is desirable to perform motion compensation or residual compression in the feature space and more effectively reduce spatial or temporal redundancy in videos. Furthermore, the recent progress in deformable convolution has shown it is feasible to align two consecutive frames in the feature space. Based on the learned offset maps of convolution kernels (i.e., the so-called dynamic kernels), deformable convolution can decide which positions are sampled from the input feature. By using deformable convolution based on dynamic kernels, we can better cope with more complex non-rigid motion patterns between two consecutive frames, which can thus improve the motion compensation results and also alleviate the burden for the subsequent residual compression module. Unfortunately, it is still a non-trivial task to seamlessly incorporate the feature space operations and deformable convolution into the existing hybrid deep video compression framework and train the whole network in an end-to-end fashion.

In our work, instead of following the existing works to adopt the pixel-level operations as in the traditional video codecs, we propose a new feature-space video coding network (referred to as FVC) by reducing the spatial-temporal redundancy in the feature space. Specifically, we first estimate motion information (i.e., the offset maps for convolution kernels in deformable convolution), based on the extracted representations from two consecutive frames. Then the offset maps are compressed by using the auto-encoder style network and the reconstructed offset maps will be used in the subsequent deformable convolution operation to generate the predicted feature for more accurate motion compensation. After that, we adopt another auto-encoder style network to compress the residual feature between the original input feature and the predicted feature. Finally, a multi-frame feature fusion module is proposed to combine a set of reference features from multiple previous frames for better frame reconstruction.

When compared with the state-of-the-art learning based video compression methods , we perform all operations in the feature space for more accurate motion estimation and motion compensation, in which we can seamlessly incorporate deformable convolution into the hybrid deep video compression framework. As a result, we can alleviate the errors introduced by inaccurate pixel-level operations like motion estimation/compensation and achieve better video compression performance. Our contributions are summarized as follows:

We propose a new learning based video compression framework, which performs all operations including motion estimation, motion compensation and residual compression in the feature space.

For effective motion compensation in the feature space, we use deformable convolution to “warp” the reference feature from the previous reconstructed frame and more accurately predict the feature of the current frame, in which the corresponding offset maps are compressed by using the auto-encoder style network.

We propose a new multi-frame feature fusion module based on the non-local attention mechanism, which combines the features from multiple previous frames for better reconstruction of the current frame.

Our framework FVC achieves the state-of-the-art performance on four benchmark datasets including HEVC, UVG, VTL and MCL-JCV, which demonstrates the effectiveness of our proposed framework.

Related Work

In the past decades, the traditional image compression methods like JPEG , JPEG2000 and BPG have been proposed to reduce spatial redundancy, in which the hand-crafted techniques, such as DCT, are exploited for achieving high compression performance. Recently, the learning based image compression methods have attracted increasing attention. The RNN-based image compression methods are firstly introduced to progressively compress images. Other methods adopted the auto-encoder structure to first encode images as latent representations, and then decode the latent representations from the feature space to the pixel space, which have achieved the state-of-the-art performance for image compression.

2 Video Compression

A number of video compression standards have been proposed in the past few years. Most methods follow the hybrid coding structure, where the motion compensation and residual coding techniques are used to reduce the redundancy in both spatial and temporal domains. Recently, learning based video compression has become a new research direction. Following the traditional hybrid video compression framework, Lu et al. proposed the first end-to-end optimized video compression framework, in which all the key components in H.264/H.265 are replaced with deep neural networks. Lin et al. used multiple frames in different modules to further remove the redundancy, while Hu et al. proposed a resolution-adaptive optical flow compression method by automatically deciding the optimal resolution for each frame and each block.

Most of the existing approaches have to estimate the pixel-level optical flow maps and compress the corresponding pixel-level residual information. However, it may be difficult to produce reliable pixel-level flow maps or residual information, which often degrades the compression performance of learning based video codecs. To address this problem, some recent works proposed to calculate the feature-level residual maps for video compression. In contrast to these works , in our work, we propose to perform all operations in the feature space, which leads to better video compression performance.

3 Deformable Convolution

Recently, Dai et al. proposed to use deformable convolution together with the learnt offset maps to enhance the modeling capability of convolution neural networks, which has achieved promising results in several video-related tasks like action recognition and video super-resolution . In contrast to these works, our work is the first to investigate how to employ deformable convolution in learning based video compression. Considering that there are multiple complex components in our hybrid video compression system, it is a non-trivial task to develop an end-to-end optimized video codec by seamlessly incorporating deformable convolution and simultaneously compressing the corresponding offset maps and other features.

Methodology

Let X={X1,X2,...,Xt−1,Xt,...}X=\{X_{1},X_{2},...,X_{t-1},X_{t},...\} denote a video sequence, where XtX_{t} is the original video frame at the current time step. In our video compression system, the objective is to produce the high quality reconstructed frame X^t\hat{X}_{t} at any given bit-rate. As shown in Fig. 1, all the modules in our proposed framework including deformable compensation, residual compression and multi-frame feature fusion are performed in the feature space. The overview architecture of our proposed framework is summarized as follows,

Feature Extraction. To produce the representations in the feature space, the original input frame XtX_{t} and the previous reconstructed frame X^t−1\hat{X}_{t-1} are encoded as the feature representations FtF_{t} and Ft−1refF^{ref}_{t-1}, respectively. As shown in Fig. 2(a), the feature extraction module uses a convolution layer with the stride of 2, which is then followed by several residual blocks to produce the feature representation.

Deformable Compensation. This procedure consists of three steps: motion estimation, motion compression and motion compensation. Specifically, based on FtF_{t} and Ft−1refF^{ref}_{t-1}, we perform motion estimation by using a lightweight network, and the output offset map OtO_{t} will be compressed by using the newly proposed motion compression network before being transmitted to the decoder side. Finally, given the reconstructed offset map O^t\hat{O}_{t} and the feature Ft−1refF_{t-1}^{ref}, we can generate the predicted feature Fˉt\bar{F}_{t} by using deformable convolution in the motion compensation procedure. More details can be found in Section 3.2.

Frame Reconstruction. As shown in Fig. 2(b), the feature decoder with several residual blocks and a deconvolution layer will transform the final reconstructed feature F^t\hat{F}_{t} to the reconstructed frame Xt^\hat{X_{t}}.

Entropy Coding. The quantized features from the motion compression and residual compression modules will be transformed into the bit-streams by performing entropy coding. During the training process, we use a bit estimation network to estimate the number of bits. More details will be described in section 3.4.

2 Deformable Compensation

The previous video compression frameworks rely on the optical flow estimation network and the pixel-level motion compensation network, which may lead to inaccurate frame prediction results and thus bring extra redundancy to the subsequent residual compression module. To this end, we perform motion compensation in the feature space and generate the predicted features by using the deformable convolution operation.

Given the features Ft−1refF^{ref}_{t-1} and FtF_{t} respectively from the previous reconstructed frame and the current frame, the deformable compensation procedure aims at generating the predicted feature Fˉt\bar{F}_{t} at the current time step. The whole network architecture is shown in Fig. 3. Specifically, we first produce the deformable offset map OtO_{t} by feeding Ft−1refF^{ref}_{t-1} and FtF_{t} into a 2-layer motion estimation network. Inspired by , we propose an auto-encoder style network to compress this offset map OtO_{t}, where both the encoder and the decoder consist of a set of Resblocks . The offset map OtO_{t} is transformed to the latent space through the encoder and then quantized. After that, the decoder will convert the quantized latent representations back to the reconstructed offset map O^t\hat{O}_{t}.

Finally, in the motion compensation procedure, the deformable convolution layer takes Ft−1refF^{ref}_{t-1} as the input and then performs the deformable convolution operation by using the corresponding filters with the aid of O^t\hat{O}_{t}. In this way, we can better compress videos with complex non-rigid motion patterns by using the learned dynamic offset kernels and produce a more accurate warped feature. To generate a more accurate predicted feature, we further refine the output feature from the deformable convolution layer by using two convolution layers and eventually produce the final predicted feature Fˉt\bar{F}_{t}.

Deformable Convolution. The network structure of our deformable convolution layer is shown in Fig. 4. Given the reference feature map and the corresponding offset map OO, we aim to generate the predicted feature. For each kernel, the corresponding offsets are used to control the sampling location in the reference feature map, and then the feature values from the corresponding locations in the reference feature map are fused to generate one feature value in the predicted feature map. When compared with the pixel-level warping operation based on motion compensation as used in the previous video compression approach , our deformable compensation module provides more flexibility by using different sampling locations, which can better cope with complex non-rigid motion patterns and improve the compensation performance. In our implementation, we divide the channels of the reference feature into G groups (G=8 in this work) and use a shared offset map for each group of channels to better learn the offset maps and improve the deformable compensation results.

3 Multi-frame Feature Fusion

4 Residual Compression and Other Details

As shown in Fig. 1, in addition to offset map compression, we also need to compress the residual feature map. To simplify the whole system, we adopt the same network architecture as that used for the offset feature map compression (see Fig. 3(b)).

For the whole learning based video compression framework, we use the bit-rate estimation network to generate the bit-rate in the training stage. In our implementation, we employ the hyperprior entropy model in for accurate bit-rate estimation. To reduce the computational complexity, the time-consuming auto-regressive model is not used in our approach.

To optimize the whole model in an end-to-end manner, it is also required to design a differentiable quantization operation. In our approach, we follow the method in and approximate the quantization operation by adding the uniform noise in the training stage. In the inference stage, we directly use the rounding operation.

5 Loss Function

In our proposed framework, we optimize the following Rate Distortion (RD) trade-off:

where RoR_{o} and RrR_{r} denote the numbers of bits used to encode the offset map OtO_{t} and the residual feature map RtR_{t}, respectively. d(Xt,X^t)d(X_{t},\hat{X}_{t}) denotes the distortion between the input frame XtX_{t} and the reconstructed frame X^t\hat{X}_{t}, where d(⋅)d(\cdot) represents the mean square error or MS-SSIM . λ\lambda is a hyper parameter used to control the rate-distortion trade-off.

Experiments

Training Dataset. We use the Vimeo-90K dataset in the training stage, which contains 89,800 video clips with each video having 7 frames with the resolution of 448×256448\times 256. We randomly crop the video sequences into the resolution of 256×256256\times 256 before the training process.

Testing Datasets. To evaluate the performance of our FVC, we use the video sequences in four datasets including the HEVC , UVG , MCL-JCV and VTL . The HEVC datasets contain 16 videos in Class B, C, D and E with different resolutions from 416×240416\times 240 to 1920×10801920\times 1080. The UVG dataset contains seven high frame rate videos with the resolution of 1920×10801920\times 1080, in which the difference between the neighboring frames is small. The MCL-JCV dataset is widely used for video quality evaluation, which consists of thirty 1080p video sequences. For the VTL dataset, we use the first 300 frames from the high-resolution video sequences with the resolution of 352×288352\times 288 for performance evaluation.

Evaluation Metrics. We use bpp (bit per pixel) to measure the average number of bits used for motion coding and residual coding for one pixel in each frame. We use PSNR and MS-SSIM to evaluate the distortion between the reconstructed frame and the ground-truth frame and the PSNR/MS-SSIM of each video sequence is produced by calculating the average PSNR/MS-SSIM over all reconstructed frames.

Implementation Details. We train four models with different λ\lambda values (λ\lambda=256, 512, 1024, 2048). A two-stage training scheme is used to train our model. In the first training stage, we train the model without using the multi-frame feature fusion module for 2,000,000 steps at the learning rate of 5e-5. In the second stage, the multi-frame feature fusion module is included and we first train the overall framework with the learning rate of 5e-5 for 400,000 steps, and then optimize the model with the reduced learning rate of 5e-6 for 100,000 steps. When using MS-SSIM for performance evaluation, we further fine-tune the model from Stage 2 for 80,000 steps by using MS-SSIM as the distortion loss to achieve better MS-SSIM results. Our method is implemented based on PyTorch with CUDA support. All the experiments are conducted on the machine with a single NVIDIA 2080TI GPU (11GB memory). We set the batch size as 4 and use the Adam optimizer . It takes about 4 days, 3 days and 12 hours for the first training stage, the second training stage and the fine-tuning stage, respectively.

2 Experimental Results

Settings of the Baseline Methods. In this section, we provide the experimental results to demonstrate the effectiveness of our proposed method FVC on four benchmark datasets HEVC , UVG , VTL and MCL-JCV . We compare our method with other state-of-the-art methods including the traditional methods and recently proposed learning based methods (DVC , AD_ICCV19 , EA_CVPR20 , LU_ECCV20 and HU_ECCV20 ). For the traditional method H.265 , we follow the command line in and use FFmpeg with the default mode. For fair comparison, we additionally report the results of “DVC*”, which is an enhanced version of DVC by using the same motion and residual encoder/decoder modules as those in our deformable compensation module (see Fig. 3). Following the previous methods , we set the GoP size as 10 for the HEVC datasets and 12 for other datasets. We use the same I-frame compression method as in H.265 to reconstruct I-frame. For our “FVC#”, we further perform the fine-tune operation according to the MS-SSIM based rate-distortion trade-off.

Results. In Table 1, we report the BDBR results of different video compression methods when compared with H.265 on the HEVC, UVG, VTL and MCL-JCV datasets. Specifically, our approach saves more than 18% bit rate in terms of the overall results on all benchmark datasets. It is obvious that our approach outperforms DVC and other recent state-of-the-art learning based video compression approaches. For example, when compared with H.265, our proposed FVC saves 18.39% bit-rate on the HEVC Class D dataset while the corresponding bit-rate savings are 8.29% and 1.77% for the recent approaches LU_ECCV20 and HU_ECCV20 .

We provide the RD curves of different compression methods in Fig. 6, it is noted that our method outperforms the baseline method DVC and the enhanced version DVC* by a large margin on all datasets. We would like to highlight that DVC* directly compresses the pixel-level optical flow maps and the pixel-level residual maps by using the same compression network as our proposed method FVC, so the performance improvement clearly demonstrates the effectiveness of our proposed method. When compared with the traditional method H.265, our method achieves better results in terms of PSNRs at all bit-rates. Our method also outperforms all other baseline methods in terms of MS-SSIM.

3 Ablation Study and Analysis

Effectiveness of Different Components. As shown in Fig. 7, we take the HEVC Class D dataset as an example to demonstrate the effectiveness of different modules in our proposed framework FVC. As shown in Fig. 7, to demonstrate the effectiveness of the non-local attention (NLA) module, we remove NLA in the multi-frame feature fusion (MFF) module and the results of FVC w/o NLA drop nearly 0.2dB when compared with our complete model (i.e., FVC). We also introduce the basic model of our approach (i.e., FVC w/o MFF or FVC-basic), where the multi-frame feature fusion module is removed in our complete model. When compared with our complete model (i.e., FVC), the result of our basic model FVC-basic after removing the MFF module drops 0.5dB at 0.3bpp (see the green and black curves). These experimental results demonstrate the effectiveness of our newly proposed MFF module equipped with the NLA mechanism for fusing the features from multiple previous frames.

To verify the effectiveness of our proposed feature-level operations, in Fig. 7 we report the results of two variants of our FVC-basic. (1) FVC-basic (FS-motion & PS-residual). The predicted feature Fˉt\bar{F}_{t} after our deformable compensation module is converted as the predicted frame by using our frame reconstruction module, which is then used as the input to the pixel-space residual compression module in DVC*. (2) FVC-basic (PS-motion & PS-residual. We perform both motion-related operations and residual compression at the pixel space, which is conceptually the same as DVC* except some subtle differences in implementation details.

It is noted that FVC-basic outperforms FVC-basic (FS-motion & PS-residual) by 0.3dB at 0.38bpp, which indicates it is beneficial to perform residual compression at the feature space. Furthermore, FVC-basic (FS-motion & PS-residual) achieves 1.2dB improvement at 0.4bpp when compared with FVC-basic (PS-motion & PS-residual), which demonstrates it is also necessary to perform motion compensation at the feature space.

Analysis of Deformable Compensation. For better comparison of the motion compensation module in the feature space and the pixel space, we also take the predicted feature Fˉt\bar{F}_{t} as the input of our frame reconstruction module to produce the predicted frame and evaluate the PSNR of the predicted frame and the corresponding bpp for compressing motion information. As shown in Fig. 8, the predicted frames of our proposed FVC achieves 1.75dB improvement at 0.017bpp on the HEVC Class C dataset when compared with the predicted frames of DVC*. It should be mentioned that DVC* uses the same compression method based on the same auto-encoder style network as our proposed method FVC, so the result clearly demonstrates the effectiveness of our proposed deformable compensation module.

Visualization of Deformable Compensation Results. In Fig. 9, we take the 6th frame in the RaceHorses sequence from the HEVC Class C dataset as an example and visualize motion information and the motion compensation results. We visualize the optical flow map used for motion compensation in DVC* (see Fig. 9(b)) and two representative offset maps (from the total number of 7272 offset maps) used for deformable convolution in our work (see Fig. 9(c) and Fig. 9(d)). It can be observed that the learned offset maps in our work encode similar motion patterns as the optical flow map from DVC*. We also observe that the motion compensation result for one patch after using our proposed deformable compensation (DC) module are visually more similar to the ground-truth patch when compared with that by using the optical flow based motion compensation method in DVC*. In practice, we can achieve 0.27dB gain by only using 65% bpp for motion coding (see Fig. 9(e), Fig. 9(f) and Fig. 9(g)).

Running Time and Model Complexity. The total number of parameters in our proposed framework is about 26M, in which the parameters from the compression network used for offset maps and residual feature maps take more than 24M. We use the videos with the resolution of 1920×10801920\times 1080 to evaluate the inference time on the machine with a single 2080TI GPU (11GB Memory). The coding time of our proposed frameworks FVC and FVC-basic as well as DVC and DVC* are 548ms, 201ms, 460ms and 709ms, and the corresponding BDBRs on the HEVC Class D dataset by using H.265 as the anchor method are -18.39%, -7.09%, 14.08% and 4.31%. From the results, we observe that our FVC-basic outperforms DVC and DVC* in terms of both efficiency and effectiveness. In addition, our proposed framework FVC achieves the best BDBR result with 347ms used for the multi-frame feature fusion module, which can be accelerated by using a simpler fusion module in our future work.

Conclusion

In this work, we have proposed a new framework FVC for deep video compression in the feature space, which consists of the deformable compensation module, the feature-level residual compression module and the multi-frame feature fusion module. The deformable compensation module first predicts the offset maps as motion information, which are then compressed by using the newly proposed motion compression module and the reconstructed offset maps will be finally used in deformable convolution for motion compensation. The proposed multi-frame feature fusion module takes multiple feature representations from the current frame and the previous frames and uses the deformable compensation and the non-local attention mechanism to refine the initial reconstructed feature for better frame reconstruction. By performing all the operations in the feature space, our framework achieves promising results on the HEVC, UVG, VTL and MCL-JCV datasets.

Acknowledgement This work was supported by the National Key Research and Development Project of China (No. 2018AAA0101900).

References