Implicit Neural Video Compression
Yunfan Zhang, Ties van Rozendaal, Johann Brehmer, Markus Nagel, Taco Cohen
Introduction
Video streaming makes up a major portion of today’s internet traffic . Compression codecs based on deep learning have recently become competitive with popular classical codecs like H.264 (AVC) and H.265 (HEVC) , but these methods have not yet been widely adapted in real-life applications. One reason for this is that neural codecs do not yet robustly outperform traditional codecs in terms of compression performance. At least as important, however, are practical considerations: neural codecs require access to a (typically large) neural network on each device on which videos need to be decompressed. This is memory-heavy, difficult to maintain, and can be vulnerable to corruption. A lightweight, computationally efficient neural codec that does not require storing large network weights might be more practical, especially for on-device applications. Moreover, standard neural codecs require a training dataset that is similar to the video samples expected at test time; the compression performance potentially suffers under training set bias and domain shift, for instance when networks trained on natural scene data are used to compress animated sequences.
We propose implicit pixel flow (IPF), a new method for video and image compression based on implicit neural representations (INR) that addresses these practical shortcomings. Each frame is represented as a function that maps coordinates within the frame to RGB values. We implement these functions as neural networks, building on recent progress in neural scene representation. Encoding then consists of choosing the architecture and overfitting the network weights on the video frames. Decoding only requires forward passes of the network. We quantize the neural network weights with fixed-point integer quantization with learned parameters and separate per-channel bit widths.
To further reduce the bitrate for video data, we compress most frames as P-frames, i. e. using the information from the previous frame. We leverage the similarity of subsequent frames through a separate implicit network that outputs the optical flow field, a map of displacement vectors that can be applied to the previous frame to approximate the current frame. We argue that implicit neural representations are a natural fit for such an optical flow warping operation: they require a simple addition in the input space, avoiding the usual interpolation-based operations that are computationally expensive and difficult to implement on device . In addition to the lightweight flow network, we train an equally lightweight residual network to complete the modeling of a P-frame.
One key advantage of our codec is that it eliminates the need to store a neural network on the receiver side, instead it only requires a framework for the evaluation of neural networks to be present. Due to the efficient flow warping operation, decoding is less computationally expensive and more hardware-friendly than established neural video codecs. Finally, this method does not require any separate training dataset. Not only does this avoid potential privacy concerns, it also means that our codec performs well on data from different domains, including those for which no suitable training data is available.
Related Work
Neural video compression typically follows the framework of variational or compressive autoencoders consisting of an encoder that maps images or video frames to a latent variable , a decoder that maps the latent variables to the reconstructed image, and a prior under which the quantized latent variables are entropy-coded. On a training dataset, these models learn to minimize the rate-distortion () loss
The trained prior and decoder need to be available at the receiver side in order to decode the transmitted bitstream. Works following this scheme include Refs. , which use 3D convolution architectures, and Refs. , which model P-frames as an optical flow field applied to the previous frame plus a residual model.
Recently, Ref. introduced instance-adaptive fine-tuning, in which the full autoencoder model is fine-tuned on each test instance. The network weight updates are entropy-coded to the bitstream and transmitted alongside the latents. While this approach relaxes the requirements on the model to generalize from the training dataset to any instance encountered at test time, it still requires a pretrained decoder and prior model to be available at the receiver side.
The first publication to apply implicit neural representation to compression is Ref. . The authors propose to compress images through their implicit representation as neural network weights. They focus on SIREN models with varying numbers of layers and channels and quantize them to 16-bit precision.
In this work, we propose a video codec that features an I-frame codec with several improvements over Ref. . Namely, we share the compute across neighboring pixels to reduce complexity and deploy a learned quantization scheme that can achieve better bitrate efficiency, and that can be trained end-to-end using a rate-distortion loss.
During the final preparations of this manuscript, two concurrent works appeared. In Ref. the authors train a single model over the entire video. For the spatial dimensions, the authors use an upsampling procedure similar to ours, achieving faster training and inference speeds. The authors then use pruning as well as quantization to reduce the bit-rate of the trained networks. The resulting RD performance is impressive, matching traditional codecs at low bitrates. However, such models require large chunks of a video to be encoded at once and cannot be applied in a low-delay setting. On the other hand, Ref. introduces an image compression codec based on quantization-aware training and meta-learned initializations. The result shows improvement over JPEG as well as our base models.
Implicit neural representations
Implicit representations have been successfully used for learning three-dimensional structures and light fields (see and references therein). Similar to our approach to compression, these works train a neural network on a single scene such that it is encoded by the network weights. New views of the scene can then be generated through a forward pass of the network. These methods are proven to be more efficient than their discrete counterparts due to their ability to exploit the redundancies.
While implicit representations have also been applied to data with lower-dimensional coordinates such as images and videos , the relative efficiency compared to discrete or latent representations has not been carefully studied. Furthermore, implicit representations have to compete with established compression codecs. Reference , to the best of our knowledge the only comparison of implicit image compression methods to classical codecs yet, demonstrated that this is no easy task.
Regardless of the dimension of input data, choosing the right class of representations is important. It has been shown that Fourier-domain features are conducive to implicit neural models learning the structure of realistic scenes. Specifically, Ref. proposed using randomly sampled Fourier frequencies as encoder prior to passing in MLP model. Ref. showed that pixel-space MLPs can achieve comparable results when using sinusoidal activations, provided that the weights are initialized carefully.
Dynamic scene representations
Implicit representations are continuous in nature. Shifting the input to these networks corresponds to continuous spatial translations within the represented scene or image. In this way, implicit representations lend themselves to tasks such as super-resolution or warping of 3D scenes . We extend this idea to motion compensation between video frames and propose to compress the displacement map with another implicit neural representation. The closest related work we are aware of is Ref. , which introduces an auxiliary network to model dynamic warping between smartphone selfies. They introduce a reference coordinate frame for each scene and compute the warping of each picture with respect to the reference. Unlike selfies, videos frames are inherently sequential. We take advantage of this natural ordering in designing our method.
Model quantization
Neural network quantization is an active field of research with the goal of reducing the size of models and running them more efficiently on resource-constrained devices. Good surveys of the field are . On a high level, there are two lines of work: vector quantization , which represents a quantized tensor using a code book, and fixed-point quantization , which represents a tensor with a fixed-point number that consists of a integer tensor and a scaling factor. In this paper we focus on the latter as it enables memory savings and leads to reduced compute complexity during training. The fixed-point quantization function is defined as
where is a integer tensor with bits and is a scaling factor (or vector) in floating point. We will use the symbol to refer to the set of all quantization parameters.
Low-bit quantization of all weight tensors in a neural network can incur significant quantization noise. With quantization-aware training , neural networks can adapt to the quantization noise by training them end-to-end with the quantization operation. As the rounding operation in Eq. (2) is non-differentiable, commonly the straight-trough estimator (STE) is used to approximate its gradient. Next to learning the scaling factors jointly with the network, recent work also started on learning a per-tensor bit-width for every layer . We extend the approach of to learn a per-channel bit-width and scaling factor. Contrary to earlier approaches, we formulate the quantization bit-width as a rate loss, and minimize the -loss (Eq. (1)) to learn the best trade-off between bitrate and distortion in pixel space.
Implicit pixel flow
Our proposed video compression codec is based on representing images through neural networks that map coordinates to RGB values. The compressed file thus consists of a small header that specifies the network architecture, and the weights of a neural network for each frame. On the one hand, an expressive architecture ensures that the network can learn the image with high fidelity. On the other hand, we quantize the network weights to keep its size down.
To compress video data more efficiently, we make use of the similarity of successive frames. We split the video in small blocks of frames (“groups of pictures” or GoP). The first frame in a GoP is compressed as an I-frame, training a single network to compress an image. The remaining frames in the GoP are trained as P-frames, i. e. using the previous frame as reference.
For P-frames, we introduce two ideas that reduce the bitrate. First, we introduce a separate implicit representation of the optical flow field or estimated motion vector, which describes how much different parts of the scene moved between frames. By adding this flow field to the inputs of our neural representation, we get a sensible first approximation of the current frame. This form of motion compensation (or optical-flow warping) is a much simpler and less computationally expensive operation than the interpolation-based approach typically used in neural video codecs . Second, we introduce a residual model that further improves the distortion.
Encoding a video consists of training a network with quantized weights, minimizing the rate-distortion loss
In the following, we will discuss in more detail our image representation in Sec. 3.2, followed by the video structure and the optical flow warping in Sec. 3.3. In Sec. 3.5 we discuss the training pipeline in more detail, before finally describing the quantization and entropy coding scheme in Sec. 3.4.
2 Implicit image representations
At the heart of our compression codec is the choice of neural network architecture used to represent individual images. Like most neural implicit models, we use models that take coordinates within an image as input and return RGB values,
The most common models used in the literature are multi-layer perceptrons (MLP). Specifically, we find that the SIREN architecture allow the highest expressivity . By using periodic activation functions, these models ensure that fine details in images and videos can be represented accurately. Decoding an image entails evaluating the MLP at every pixel location of interest.
The input space of this representation is continuous. We are thus not tied to any particular pixel grid, and the representation can be trained or evaluated at different resolution settings or on irregular grids. Our method is therefore suitable for super-resolution tasks or for decoding video data directly in a target resolution depending for instance on the available hardware.
Upsampling SIREN (uSIREN)
While a SIREN-based MLP is expressive, it requires one forward pass for each pixel in decoding an image, which can quickly get expensive on full resolution media. To lighten the decoding compute, we share part of the computation between neighboring pixels. To see how this works, notice that a MLP can also be regarded as a convolution, where the convolving dimensions are the coordinate dimensions of the input. By inserting a bilinear interpolation layer of stride between the convolution layers, we reduce the operating dimension or layers preceding the upsampling by a factor of . This means the decoding compute of those layers are also reduced by a factor of .
3 Implicit video representations
Video data often have strong redundancies between subsequent frames. Neural implicit representations can represent video data by extending the input space by a third time or frame dimension . While this approach is straightforward, we find that the implicit networks are not expressive enough to represent high-resolution video data at low distortion, as we demonstrate in experiments in the supplementary materials. In addition, this 3D approach in less suitable for video streaming, because an entire block of frames needs to be received before even the first frame can be decoded.
Instead, we propose to compress video sequences frame by frame while still leveraging the similarity between them. Like many classical and neural video codecs, we split the full video sequence into multiple GoPs. In each GoP, only the very first frame is compressed as a stand-alone I-frame, while all other frames are compressed using available information from other frames. Here we focus on the common case where each frame in a GoP except the first is compressed as a P-frame: it explicitly depends on the previous frame, but not future frames. This setup is well-suited to video streaming in a low delay setting.
One popular strategy to compress P-frames is to estimate the optical flow field or motion vector, a field of displacement vectors that, applied to the previous frame, give a good first estimate of the current frame. In addition to the optical flow, pixel-level or feature-level residuals are compressed. The final prediction for the frame is given by applying the optical flow field to the previous reconstructed frame in an operation known as warping or motion compensation and adding the residuals. While this approach can substantially lower the bitrate, optical-flow warping requires expensive and hardware-unfriendly interpolation operations.
Implicit flow warping
In this work we model optical flow implicitly by leveraging the fact that implicit representations are continuous. Recall that frames are represented as a network that takes image coordinates as input, . Applying the displacement from an optical flow field requires only to add the displacement vector to the input variables:
The displacement fields are represented implicitly as neural networks with weights , using smaller SIREN architectures (Sec. 3.2).
Residual modeling
Even an accurate flow model cannot always fully predict a frame, for instance due to occlusion effects or the introduction of new objects. We therefore model residuals on top of the warped frame,
with a separate implicit network . The same small model as for the flow suffices for a good performance, allowing us to keep the bitrate low. In ablation studies we found this modeling of residuals in pixel space to be more effective than modeling residuals in the weight space of the implicit networks.
Sometimes, implicit flow is all you need
Like many autoencoder-based video codecs, our IPF model thus follows the basic recipe of modelling P-frames as the previous frame warped by a compressed flow field plus a compressed residual. However, the learning dynamics lead to some practical differences: unlike in the autoencoder approach, the bitrate spent on flow fields is largely determined by the model architecture, which is fixed prior to training. Therefore there is less of an incentive for the flow field to be smooth and to consist of large near-constant patches. Instead, the IPF flow can model changes in fine-grained detail compared to the previous frame. In our experiments, this works so well that for many frame transitions the residual model is redundant.We find this behaviour independent of whether we train flow and residual model jointly or sequentially. Of course, this does not work when pixels in new colors appear, for instance because of new objects entering the scene.
We embrace this expressiveness of the flow and explore a optimized codec where the residual model is optional. In this case we dynamically decide on the inclusion of a residual model on a per GoP basis. On the encoder side, for each GoP we compute the rate and distortion performance both with and without the residual model. We choose the variant that leads to a lower loss and signal this choice by transmitting a single bitstream. If flow without residuals is the better choice for a frame, we do not need to transmit the residual model, reducing the bitrate.
4 Quantization and entropy coding
To reduce the model size of the implicit models representing I-frames, optical flow, and residuals, we quantize every weight tensor using fixed-point representation (cf. Eq. (2)). To learn the quantization parameters and bit-width jointly with the model weights, we follow the parameterization suggested by Ref. and learn the scale and the clipping threshold . The bit-width is then implicitly defined as
Reference showed that this parameterization is favorable over learning the bit-width directly as it does not suffer from an unbounded gradient norm. We further extend this approach to per-channel quantization allowing us to learn a separate range and bit-width for every row in the matrixOutput channel in case of a convolutional layer.. Our per-channel mixed precision quantization function is defined as:
Next, we encode all quantization parameters and all integer tensors to the bitstream. The are encoded as 32-bit floating point vectors, the bit-widths as 5-bit integer vectors, and the in their respective per-channel bit-width .
5 Training pipeline
We summarize our architecture in Fig. 1 and detail the training procedure in Alg. 1. Initially, unquantized networks are pre-trained by minimizing distortion, then the quantizers are added and the quantized models are trained on the combined rate-distortion loss of Eq. (3). The quantized weights are written to the bitstream.
We train the base model for each I-frame, and flow and residual models for each P-frame. While each residual model is independent, we find it beneficial to accumulate the flow across frames inside a GOP. This way each flow model only needs to model the motion between consecutive frames, instead of directly from the I-frame.
Unrolling the recursive definition of warping, the implicit representation of the P-frame is given by
where for readability we leave out the quantizer and is the implicit parameterization of the previous I-frame. In the second line, we show that we do not need to store the previous flow networks in memory and evaluate them again for each frame; instead, we just need to store the cumulative flow field in a single tensor. This tensor can be constructed in the same way on receiver side. Training the flow and residual networks for a P-frame then just amounts to minimizing the loss in Eq. (3) with the parameterization in Eq. (9).
Experiments
We now demonstrate our compression codec in experiments. First we demonstrate the ability of our neural implicit model to represent images, before turning to video data.
In Fig. 2 we show an image from the CLIC 2020 challenge, a version compressed with JPEG, and a version compressed with our neural implicit base model (see the supplementary materials for a detailed description of the setup). Both codecs are run at a strong compression setting, the image with resolution is in both cases compressed to a 200 kB file. We confirm that the network can represent the image including fine-grained detail. Compared to JPEG, it achieves a substantially better PSNR at the same file size. At this low-filesize setting, there are certainly artifacts, but they are less pronounced than those induced by JPEG.
In addition, we compare to an unquantized implicit representation with network weights stored at 32-bit floating-point precision. Quantization is able to reduce the network size substantially at a small cost in the distortion performance. Note that we are able to quantize the parameters to an average of 9.3 bits per parameter without a substantial drop in distortion performance.
In Fig. 3 we evaluate the image compression performance on the Kodak image dataset . The rate-distortion results in Fig. 3 show that our method outperforms JPEG at low to medium bitrates, but is not competitive with state-of-the art codecs like BPG. We also compare to COIN , which uses very similar implicit networks, but a simpler weight compression scheme. We find that our learned integer quantization, which allows us to compress the networks to around 9 to 10 bits per parameter, leads to a substantial performance improvement over COIN with its 16-bit floating point quantization.
In Fig. 4 we show the effect of different quantization strategies on rate-distortion performance. Floating point quantization, the most naive quantization baseline, typically only supports to 16 bits/parameter and is outperformed by the other two quantization methods which can give similar distortion at 10–11 bits/parameter. For all bitrates, learned per-channel quantization outperforms the fixed bitwidth quantization. The learned-bitwidth quantization which we use in our main models can quantize up to 8–9 bits/parameter on average with only a neglegible drop in distortion performance. This quantization strategy allows the model to allocate a different bitwidth for each channel in every parameter, as required for the best rate-distortion tradeoff. This effect is illustrated in Fig. 7. This plot shows the distribution of learned bitwidths for each parameter in the network. Especially the first and the last layers of the network require quantization to higher bitwidths, while the other layers are quantized to around 8 bits/parameter.
Video compression
We compress the 7 videos from the UVG-1k dataset , which have a Full-HD resolution ( pixels). As the content and style of these sequences do not change over the video, we only use the first half of each sequence (i. e. the first 300 frames) to save compute resources. We use three different architectures, corresponding to different working points on the rate-distortion curve. We specify the I-frame codec sizes to roughly cover the range of bit-rates we are interested in. For each model, the flow and residual models are then specified to be the size of the I-frame codec, a ratio we empirically determined to give good rate-distortion performance. We describe further details of the setup in the supplementary materials.
As baselines we compare to the popular classical codecs mpeg4-2 , H.264 , and H.265 in the ffmpeg implementation . For our method and all baselines we use a GoP size of 5 frames and operate in the low-delay setting with only I-frames and P-frames. We trained our pipelines on TeslaV100 GPUs for approximately 300 GPU hours for each video.We show the average performance over all videos in Fig. 5. Our method is able to compete with mpeg4-2 at low bit rates, but it is still clearly behind H.264 and H.265.
To gain some insights into this result, we breakdown the rate-distortion performance per type of frame in Fig. 6. The red solid line indicates our average performance as in Fig. 5. This is averaged over I-frames (dot dashed dark blue) and P frames (dashed light blue). The I-frame are 2 - 5 dB higher in PSNR than the P-frames, depending on the model size, while being 20 times as expensive to compress (considering the combined rate of flow and residual, each of which is 1/40 size of the I-frame).
The lower PSNR for P-frames can be partially attributed our use of much smaller models to model the P-frames. Furthermore, as discussed in Sec. 3.3, our flow model is tasked with reproducing the P-frame single handedly. Thus it becomes a question whether the residual model is always needed. Thus we show the alternate scenario where we do not transmit the residual model both for overall (yellow), and P-frame only (cyan). In the inset we can see that the residual model roughly doubles the rate of the P-frames with a slight improvement in PSNR. While for the P-frames this means a sub-par distortion result, the overall performance with residual (red) still improves upon without (yellow).
Finally, to take the best possible result regarding flow and residual, we dynamically decide whether to include the residual model in the bitstream as follows. For each GoP of each video, we compute the loss with and without the residual model, and decide whether to include residual model for the given GoP based on the comparison. Although this breaks the low-delay setting for encoding, it still allows low-delay decoding. The decision of whether to include the residual model adds only a single bit per GoP to the bit stream. The result of this dynamic approach is shown in dot dashed red line, which we see slightly outperforms codecs with and without residual models.
Conclusion
While neural compression algorithms are beginning to outperform classical codecs, they suffer from severe practical disadvantages. Not only do they require training datasets that match the data expected at test time, but they also require large pretrained neural networks on the receiver side. While such networks greatly aids rate-distortion performance, it presents an obstacle to practical use, especially on device.
We propose an image and video compression method based on implicit neural representations that avoids these obstacles. Each image is compressed by training a neural networks to learn its pixel content, quantizing the network weights, and entropy-coding them to the bitstream. For videos we introduced a frame-by-frame scheme that leverages the continuous nature of implicit representations to perform motion compensation without requiring hardware-unfriendly operations like interpolation.
While our method is not yet competitive with the rate-distortion performance of state-of-the-art codecs, we demonstrate that it can nevertheless compress image and video data efficiently. Given its practical advantages, we believe it can serve as a step towards self-contained learning-based codecs that are deployable in real life use cases.
References
Appendix A Video demonstration
We include reconstructions of the “Bosphorus” sequence from the UVG dataset for all three IPF models in the supplementary files. The videos are 300 frames at 120fps.
Appendix B Detailed results
In Fig. 9 we show breakdown of our method against baselines for each video in the UVG dataset. Similar to overall result in the main text, we show the full model with residuals, no residuals, and a model with optional residual components. We see clear difference between much static videos such as Honey Bee and more dynamic videos such as Jockey. Our method performs better compared to the baselines on the static videos than on dynamic videos. Residual component leads to more improvement on dynamic videos.
B.2 Quantization details
Here we present the learned bitwidths for the implicit model. Figure 7 shows the learned bitwidths for a single model (the model used to compress the image in Fig. 2). It can be seen that the model learns an efficient bit-width of 8.1 bits/parameter. Furthermore, we can see that the model dynamically learns to allocate bit-width, spending more bits on small parameters such as biases, or on “important” layers such as the first and the last layer.
B.3 Time implicit SIREN (SIREN3D)
When generalizing implicit neural compression models from two-dimensional images to video data, a natural choice is to incorporate the time axis as a new implicit dimension. The resulting architecture, which we here refer to as SIREN3D, operates directly on an entire group of pictures.
In Fig. 8 we compare the average performance of this model on the first 5 GoPs, each consisting of 5 frames, to that of our IPF model. SIREN3D underperforms IPF in all rate ranges tested.
The careful reader may notice the rate differences between models. The initial SIREN3D models are roughly the same size as that of the IPF I-frame models (see Sec. C for details). However, our quantization procedure uses the same value of in its optimization objective (Eq. (3)), thus the resulting rates are somewhat larger than that of their IPF counterparts.
Appendix C Implementation details
The architectures of the implicit models used in our image compression experiments are summarized in Tbl. 1. Our choice for the Kodak dataset is based on the hyperparameters described in Ref. , except for two differences. Firstly, on the very last layer we use ReLU activations while uses identity/no activations. Secondly, Ref. use networks with one more layer than they describe in their appendix; we follow their description in the paper, not their implementation, and thus use one layer less in every model. We find that our models are comparable in distortion performance to theirs even with one fewer layer. In addition, we add two larger models aimed at a higher-quality setting.
Training hyperparameters
We first train the implicit models with Adam for 100 000 steps with a learning rate decaying exponentially from to . Then we train the quantized models for an additional steps, using a constant learning rate of and weighting the bits per parameters in the loss function with for the lower-bitrate models and for the higher-bitrate models.
C.2 Video compression
The architecture hyper-parameters of our I-frame SIREN, and uSIREN (see Sec. 3.2 in the main text), as well as an alternative video model SIREN3D (see Sec. B.3) are shown in Tbl. 2. For a fair comparison, the architectures are chosen with a similar number of learnable layers. The number of channels are then chosen so that the total number of parameters match. Three model sizes (small, medium, and large) are explored. For simplicity we keep the number of layers the same across this size range. Not shown are the architectures of flow models. For the medium and large IPF, we use as flow model a 6 layer siren with 32 channels. For the small IPF models, we use 6 layers with 24 channels.
Training hyperparameters
In this section we detail some of our training setup for the full IPF. For all stages we optimize towards a rate-distortion objective based on the MSE as distortion metric and with . We use the Adam optimizer throughout. Table 3 details the learning rate schedules.
We make our training more efficient by initializing each I-frame model except for the very first at the implicit model representing the previous I-frame. With this initialization, we only need to train for 80k steps.
Similarly, all flow models (except for the first in each GoP) are initialized to the flow model from the previous P-frame. We then skip the unquantized pretraining phase and directly optimize with quantization.
Appendix D Baselines
In our first image compression analysis on an image from the CLIC 2020 Challenge (see Fig. 2, we compare to a JPEG baseline generated with Pillow with subsampling disabled.
Kodak
On the Kodak image dataset, we compare to COIN, JPEG, JPEG2000, and BPG baselines as reported by Ref. .
Video
We compare to three popular classical codecs (mpeg4-part2, H.264, H.265) in their ffmpeg (x265, x266) implementations . All methods are restricted to only I-frames and P-frames and a fixed GoP size of 5. We use the ffmpeg “medium” encoder preset.