CompressAI: a PyTorch library and evaluation platform for end-to-end compression research
Jean Bégaint, Fabien Racapé, Simon Feltman, Akshay Pushparaja
Introduction
Following the success of deep neural networks in out-performing most conventional approaches in computer vision applications, Artificial Neural Network (ANN) based codecs have recently demonstrated impressive results for compressing images .
Conventional lossy image compression methods like JPEG , JPEG2000 , HEVC or AV1 or VVC have iteratively improved on a similar coding scheme: partition images into blocks of pixels, use transform domain to decorrelate spatial frequencies with linear transforms (e.g.: DCT or DWT), perform some predictions based on neighboring values, quantize the transformed coefficients, and finally encode the quantized values and the prediction side-information into a bit-stream with an efficient entropy coder (e.g.: CABAC ). On the other hand, ANN-based codecs mostly rely on learned analysis and synthesis non-linear transforms. Pixel values are mapped to a latent representation via an analysis transform, the latent is then quantized and (lossless-ly) entropy coded. Similarly, the decoder consists of an approximate inverse transform, or synthesis transform, than converts the latent representation back to the pixel domain.
By learning complex non-linear transforms based on convolutional neural networks (CNN), ANN-based codecs are able to match or outperforms conventional approaches. The training objective is to minimize the estimated length of the bitstream while keeping the distortion of the reconstructed image low, compared to the original content. The distortion can be easily measured with objective or perceptual metrics like the MSE (mean squared error) or the MS-SSIM (multi-scale structural similarity) . Minimizing the bit-stream size requires to learn shared probability models between the encoder and decoder (priors) , as well as using relaxation methods to approximate the non-differentiable quantization of the latent values. The complete encoding-decoding pipeline can be trained end-to-end with any differentiable distortion metrics, which is especially appealing for perceptual metrics (learned or approximated) or machine-tasks related metrics (for example image segmentation/classification at very low-bitrates).
Similarly, significant progresses have been reported regarding neural networks targeting video compression . Compressing videos is more challenging as reducing temporal redundancies requires to estimate motion information (such as optical flow) involving larger networks and multiple-stages training pipelines .
Promising results have been achieved with ANN-based codecs for image/video compression and further improvements can be excepted as better entropy models, training setups and network architectures are discovered. As such, more research and experiments are required to improve the performances of learned codecs. However as this field is relatively new, there is a lack of tooling to facilitate researchers contributions. The CompressAI platform, presented in this document, aims to help improve this situation.
Motivation
The current deep learning ecosystem is mostly dominated by two frameworks: PyTorch and TensorFlow . Discussing the merits, advantages and particulars of one framework over the other is beyond the scope of this document. However, there is evidence that PyTorch has seen a major growth in the academic and industrial research circles over the last years. On the other hand, building end-to-end architectures for image and video compression from scratch in PyTorch requires a lot of re-implementation work, as PyTorch does not ship with any custom operations required for compression (such as entropy bottlenecks or entropy coding tools). These required components are also mostly absent from the current PyTorch ecosystem, whereas the TensorFlow framework has an official library for learned data compression https://github.com/tensorflow/compression/.
CompressAI aims to implement the most common operations needed to build deep neural network architectures for data compression in PyTorch, and to provide evaluation tools to compare learned methods with traditional codecs. CompressAI re-implements models from the state-of-the-art on learned image compression. Pre-trained weights, learned using the Vimeo-90K training dataset , are included for multiple bit-rate points and quality metrics, which achieve similar performances to reported numbers in the original papers. A complete research pipeline, from training to performance evaluation against other learned and conventional codecs, is made possible with CompressAI.
Design
CompressAI tries to adhere to the original design principles of PyTorch: be pythonic, put researchers first, provide pragmatic performance and worse is better . As a PyTorch library, it also follows the code structures and conventions which can be found in widely used PyTorch libraries (such as TorchVision https://github.com/pytorch/vision or Captum https://github.com/pytorch/captum).
As a research library, CompressAI aims to follow the naming conventions introduced in the literature on learned data compression, as to ease the transition from paper to code. High level APIs to train or run inference on models do not require much prior knowledge on learned compression or deep learning. However, specific code implementations relating to the (learned) compression domain may require to be familiar with the terminologies. This facilitates adoption and ease of use for researchers.
Features
One of the most important features of CompressAI is the ability to easily implement deep neural networks for end-to-end compression. Several domain-specific layers, operations and modules have been implemented on top of PyTorch, such as entropy models, quantization operations, color transforms.
Multiple architectures from the state-of-the-art on learned image compression have been re-implemented in PyTorch with the domain specific modules and layers provided by CompressAI. See Table 1 for a complete list and description. With a few dozen lines of python code, a fully end-to-end network architecture can be defined as easily as any PyTorch model (cf. Appendix C for some code examples).
Currently, CompressAI tools and documentation mostly focus on learned image compression and will soon add support for video compression. However, other end-to-end compression pipelines could be built using CompressAI, like the compression of 3D maps or deep features for example.
2 Model zoo
CompressAI provides pre-trained weights for multiple state-of-the-art compression networks. These models are available in a model zoo and can be directly downloaded from the CompressAI API. Theses pre-trained weights allow very near reproduction of the results from the original publications, as well as fine-tuning of models with different metrics or bootstrapping more complex models.
Instantiating a model from pre-trained weights is fast and straightforward (see example 1). Internally, CompressAI leverages the PyTorch API to download and cache the serialized object models.
As is common in the PyTorch ecosystem, pre-trained models expect input batches of RGB image tensors of shape (N, 3, H, W), with N being the batch size and (H, W) the spatial dimensions of the images. The input image data range should be , and no prior normalization is to be performed.
Due to the design of the reference models, some constraints need to be respected: H and W are expected to be at least 64 pixels long. Based on the number of strided convolutions and deconvolutions for a particular model, users might have to pad H and W of the input tensors to the adequate dimensions.
Most of the models have different behaviors for their training or evaluation modes. For example, quantization operations may be performed differently: uniform noise is usually added to the latent tensor during training, while rounding is used at the inference stage. Users can switch between modes via model.train() or model.eval().
Unless specified otherwise, the provided pre-trained networks were trained for 4-5M steps on image patches randomly extracted and cropped from the Vimeo-90K dataset .
Models were trained with a batch size of 16 or 32, and an initial learning rate of 1e-4 for approximately 1-2M steps. The learning rate is then divided by 2 whenever the evaluation loss reaches a plateau (we use a patience of 20 epochs). Training usually takes between 4 or 10 days to reach state-of-the-art performances, depending on the model architecture, the number of channels and the GPU architecture used.
The loss functions and parameters used for training are respectively reported in tables 3 and 2. The number of channels in the auto-encoder bottleneck varies depending on the targeted bit-rates. The bottleneck needs to be larger for higher bit-rates. For low bit-rates, below 0.5 bpp, the literature usually recommends using 192 channels for the entropy bottleneck, and 320 channels for higher bit-rates. CompressAI provides downloadable weights for most of the pre-defined architectures, pre-trained using either MSE or MS-SSIM (multi-scale structural similarity), for multiple bit-rates (6 or 8) up to 2bpp.
3 Utilities
CompressAI ships with some command line utilities that may come in handy when developing or evaluating learned image compression codecs. The following tasks can be performed directly from the command line by calling CompressAI scripts:
evaluating a pre-trained or user-trained model on a dataset of images
evaluating a conventional codec on a dataset of images
finding the right quality parameter to reach a given PSNR or bit-rate on a target image (see example in listing 2)
4 Benchmarking
This section exposes the evaluation tools implemented in CompressAI. One of the design goals was to provide simple and efficient tools for comparing end-to-end methods and traditional codecs. This allows researchers to reproduce and validate published results, but also to iterate rapidly over research ideas.
Currently, supported quality metrics are the PSNR (peak signal-to-noise ratio) and the MS-SSIM.
Rate-distortion performances of learned models can be measured for each of the included models in CompressAI, see Table 1 for an exhaustive list of the re-implemented models.
4.2 Traditional codecs
To facilitate the comparison with traditional codecs, CompressAI includes a simple python API and command line interface. The most common image and video codecs are supported, the complete list of supported codecs and their respective implementations can be found in Table 4.
Default parameters have been chosen to provide fair and comparable results between conventional codecs and learned ones.
At this time, runtime comparisons between traditional methods and ANN-based methods cannot be reported in an accurate and fair manner. Measuring the inference efficiency of an ANN-based codec is an active research topic http://www.compression.cc/. Neural networks are effectively designed to run on massively parallel architectures whereas traditional codecs are typically designed to run on a single CPU core.
CompressAI does not ship the binaries of the above traditional codecs but rather provides a common Python interface over the executables, with the exceptions of JPEG and WebP which are linked by default using the Python Pillow Image library.
Evaluation
This section exposes the equivalence between the results produced using CompressAI by retraining state-of-the-art methods from scratch using Vimeo-90K as training set, and the originally published results.
The following models have been re-implemented in CompressAI:
factorized-prior and hyperprior models from Ballé et al. ,
hyperprior with non-zero Gaussian means and auto-regressive models from Minnen et al. .
anchor and self-attention models joint hyperprior models from Cheng et al. .
The pre-trained weights, optimized for the MSE (Mean-Square-Error) metric, can be directly accessed from the CompressAI API. Pre-trained weights optimized for the MS-SSIM metric are also being added.
The following graphs in Figure 3, which report the average performance on the Kodak dataset , show that similar results have been reproduced to those reported in the original publications. Note that these are actual bit-rates counted on the produced bit-streams, not the estimated entropy values provided by the networks. The full encoding/decoding pipeline is implemented within CompressAI. However, due to floating point operations (at the auto-encoder network and probability estimation levels), reproducibility across different systems or platforms is not yet achieved. Some publications on the subject have already proposed solutions (e.g.: integer models in []), and will be considered in a future version.
2 Learned codecs versus traditional video codecs
In this section, we provide a brief objective comparison between learned codecs and traditional methods. Pre-trained models provided with CompressAI are compared with HEVC (HM version 16.20), VVC (VTM version 9.1) and AV1 (version 2.0). HEVC, VVC and AV1 are configured following their default intra mode configuration and with 8-bit YCbCR 4:4:4 inputs/outputs.
As can be seen in Figure 4, recent works on learned image compression compare favorably with already published standards such as H.265/HEVC and AV1. The most performing methods are competitive with the latest ITU/ISO codec H.266/VVC in PSNR at low bit-rates. However, learned compression frameworks can be directly optimized for complex objective metrics as long as the metric is differentiable or has a differentiable approximation. This constitutes a major asset compared to traditional hybrid codecs where it is difficult to define good strategies for block-based encoder decisions. CompressAI will soon support different metrics for training and evaluation. See Appendix A for more results with the MS-SSIM metric.
Besides, learned image and video compression codecs have only started to achieve competitive results these last 5 years. Considering the success of learned methods in other image processing and computer visions domains, significant improvements in compression performance and inference speed can be expected as more research will be performed on the subject.
3 Adoption
The first public version of CompressAI was released on GitHub in early June 2020. Since then, we have already noticed adoption from both the industrial and academic research communities. Multiple MPEG contributions in the Deep Neural Network for Video Coding (DNNVC) and Video Coding for Machine (VCM) working groups http://wg11.sc29.org/ were based on CompressAI. Several academic research groups have also started using CompressAI for their research.
Conclusion and future work
CompressAI currently implements networks for still picture coding and provides pre-trained weights and tools to compare state-of-the-art models with traditional image codecs. It reproduces results from the literature and allows researchers, developers and enthusiasts to train and evaluate their own neural-network-based codec.
Several extensions to CompressAI are planned. In the next releases, CompressAI will include additional models from the literature on learned image compression, and more pre-trained weights for perceptual metrics (e.g.: MS-SSIM ). One critical extension is to add support for video compression. Evaluation for low-delay and random-access video coding with traditional codecs, and end-to-end networks with compressible motion information modules will be introduced in the next releases. A better compatibility with TorchScript and ONNX is also being considered.
The platform is made available to the research and open-source communities under the Apache 2.0 license. We plan to continue supporting and extending CompressAI openly on GitHub, and we welcome feedback, questions and contributions.
Acknowledgements
The authors would like to thank Chamain Hewa Gamage for thoughtful discussions and valuable comments on CompressAI. The authors would also like to thank the authors of the TensorFlow Compression library https://github.com/tensorflow/compression/ for open-sourcing their code.
References
Appendix A Rate-distortion curves
A.2 MS-SSIM on Kodak
A.3 PSNR on CLIC Mobile (2020)
A.4 MS-SSIM on CLIC Mobile (2020)
A.5 PSNR on CLIC Pro (2020)
A.6 MS-SSIM on CLIC Pro (2020)
Appendix B Example Images
B.2 Kodak 20
B.3 Saint Malo
Appendix C Code examples
To demonstrate the simplicity of the approach, we include a sample code to build a simple auto encoder network in Listing 14.
The code runs with Python 3.6+, PyTorch 1.5+, Torchvision 0.5 and CompressAI 1.0+. This example model is similar to the fully factorized model presented in , which is a fully convolutional network with an entropy bottleneck , and can be replicated in a few dozens line of codes.
Note that this does not include the code to actually train the network. We provide an example training code in the examples folder on the CompressAI GitHub repository https://github.com/InterDigitalInc/CompressAI/blob/master/examples/train.py. The full training code is around 300 lines of code, with additional features such as logging and models check-pointing.