Hierarchical Conditional Flow: A Unified Framework for Image Super-Resolution and Image Rescaling
Jingyun Liang, Andreas Lugmayr, Kai Zhang, Martin Danelljan, Luc Van Gool, Radu Timofte
Introduction
Normalizing flows are powerful deep generative probabilistic models that allow for efficient and exact likelihood calculation and sampling. They have been used in the generation of image , blur kernel , and audio data. Recently, in the low-level vision community, normalizing flows have attracted much interest and have achieved promising progress for image super-resolution (SR) and image rescaling .
SRFlow is a seminal flow-based model for image SR. Unlike previous CNN-based models that learn a deterministic mapping from the low-resolution (LR) image to the high-resolution (HR) image, SRFlow learns the distribution of HR images and is able to generate diverse photo-realistic HR images. However, as shown in Fig. LABEL:fig:intro_srflow, it treats the LR image as an external conditional prior and thus is not fully invertible between HR and LR image pairs, making it hard to be used for image rescaling. Another work IRN employs an invertible neural network to learn downscaling and upscaling for image rescaling. Since the model is bijective, it can recover the input HR image with high accuracy after downscaling. Nevertheless, as shown in Fig. LABEL:fig:intro_irn, it assumes the high-frequency and low-frequency components of the image are independent to each other and thus lacks the ability to exploit their dependency for image SR.
In this paper, we propose a hierarchical conditional flow (HCFlow) as a unified framework for both image SR and rescaling. As shown in Fig. LABEL:fig:intro_mcflow, HCFlow is an invertible flow-based model for modelling the HR-LR relationship, in which the high-frequency component is hierarchically conditional on the low-frequency component of the image. More specifically, in the forward propagation, HCFlow learns to decompose the input HR image into the LR image and a latent variable. In the inverse propagation, it generates HR images based on the LR input and random samples of the latent variable. The modelling of the latent variable (high-frequency component) is conditional on the generated LR image (low-frequency component) in a hierarchical manner.
When trained for image SR, HCFlow is optimized by minimizing the negative log-likelihood loss on the basis of tractable Jacobian determinant computation. To further improve visual quality, we integrate a pixel loss, perceptual loss, and GAN loss in the inverse propagation to constrain the learned HR space. Moreover, HCFlow can be used for the image rescaling task. It can decompose the HR image to a visually-pleasing LR image and a latent variable that follows a simple distribution. In this case, HCFlow is trained as an encoder-decoder framework, in which the forward and inverse processes are jointly optimized. As HCFlow is bijective, it can recover the HR image faithfully by sampling from the latent space given the generated LR image.
Our contributions can be summarized as follows:
We propose a unified framework for image SR and image rescaling. It learns to model the LR image and the residual high-frequency component simultaneously. The high-frequency component is hierarchically conditional on the generated LR image.
We propose additional losses to train normalizing flows, including pixel, perceptual, and GAN losses, which effectively enhances the HR image quality.
We perform extensive experiments on three tasks: general image SR, face image SR and image rescaling. HCFlow achieves state-of-the-art results on all tasks in terms of both quantitative metrics and visual quality.
Related Work
In this section, we will briefly review image SR and image rescaling with a particular focus on two highly related flow-based methods, i.e., SRFlow and IRN .
Image SR aims to reconstruct the HR image given the LR image. Since the pioneer work SRCNN , many CNN-based models have been proposed in recent years . Most of them focus on delicate feature extraction module design and generate over-smoothed images when trained with the pixel loss. To remedy this, the perceptual loss and GAN loss are introduced to improve the perceptual quality. Despite of above progresses, they usually learn a deterministic mapping between the LR image and HR image, which is unnatural for image SR since one LR image may correspond to multiple HR images.
SRFlow . Normalizing flows provide a new possible solution for image SR. SRFlow designs a conditional flow to model the distribution of HR images, conditional on LR images. It can generate diverse photo-realistic images by sampling different latent variables. Our proposed HCFlow differs from SRFlow in two main aspects: First, SRFlow uses the LR image as an external conditional prior and maps the HR distribution to a simple latent distribution. Therefore, it cannot generate LR image and thus is not applicable for image rescaling. In contrast, HCFlow models the LR image and treats it as part of the latent space. Second, SRFlow basically follows the flow framework proposed in , while HCFlow proposes a new framework with hierarchical conditional mechanism.
2 Image Rescaling
Image rescaling aims to downscale the HR image to a visually meaningful LR image, and then recover the HR image plausibly. Different from image SR that works on a given LR image space, image rescaling tries to maintain as much information from the HR image as possible for a better subsequent reconstruction, for the purpose of reducing the storage and bandwidth cost. In other words, it can define its own LR image space which is expected to be more informative than that by simple downscaling such as bicubic downscaling. In general, in image rescaling, the downscaling and upscaling processes are jointly modelled by an encoder-decoder framework , so that the downscaling model is optimized for the later upscaling operation.
IRN . Recently, IRN proposes to use a bijective invertible neural network to model the downscaling and upscaling processes. High-frequency component is well-captured and transformed to a structured latent space in training. In testing, the HR image can be recovered by inputting the generated LR image and a randomly sampled latent variable. In particular, IRN assumes the LR image and high-frequency component is independent to each other. These two components are divided apart and learned separately. By contrast, HCFlow assumes the removed high-frequency component is dependent on the LR image and thus employs a hierarchical conditional framework to model the LR image and the conditional distribution of the high-frequency component. Besides, although IRN designs a bijective mapping between HR and LR image pairs, it can only be trained by Monte Carlo simulation rather than maximum likelihood estimation (MLE). HCFlow can be trained in the same way for image rescaling, but it further models the LR image distribution and allows for tractable Jacobian determinant computation, making it possible for probabilistic modelling of HR and LR images when trained by MLE.
Methodology
Flow-based models aim to learn a bijective mapping between the target space and the latent space. For a high-dimensional random variable (e.g., an image) with distribution and a latent variable with simple tractable distribution (e.g., multivariate Gaussian distribution), flow models generally use an invertible neural network to transform to : . Conversely, can be recovered from by the inverse mapping .
Generally, is composed of a series of invertible transformations: f_{\bm{\theta}}=f_{\bm{\theta}}^{1}\mathop{\scalebox{0.6}{\circ}}f_{\bm{\theta}}^{2}\mathop{\scalebox{0.6}{\circ}}\cdots\mathop{\scalebox{0.6}{\circ}}f_{\bm{\theta}}^{K}. The intermediate variables are defined as for . The input and output of are and , respectively. Concretely, are flow layers such as squeeze layer, batch normalization layer, affine coupling layer, etc.
According to the change of variable formula and the chain rule, for a sample , the log probability can be calculated as
where is the logarithm of the absolute value of the determinant of the Jacobian of at . The flow model can thereby be optimized by minimizing the negative log-likelihood loss.
2 Model Specification
Both image SR and image rescaling try to reconstruct the HR image given a LR image. Since the image degradation process (or image downscaling) is the inverse of image super-resolution (or image upscaling), we can model these two processes with an invertible bijective transformation: , where and are the generated LR image and the rest high-frequency component, respectively. As modelling the probability of natural images is a non-trivial task, it is reasonable to design a flow model conditional on the ground-truth LR image as,
Ideally, we hope the model can generate exactly the same LR image as the ground-truth LR image. This can be formulated as a Dirac delta function and further approximated by a multivariate Gaussian distribution as,
where is a diagonal covariance matrix with all diagonal elements close to zero. Note that y is nearly equal to in this case. By further mapping to a standard multivariate Gaussian distribution , the flow model is defined as,
As we can see, part of the latent space is constrained to be the LR image space. In particular, decomposed high-frequency component is conditional on another decomposed component . Once trained, following the forward direction, HCFlow can decompose the HR image into LR image and latent variable that follows a simple distribution. Following the inverse direction, HCFlow can generate given the LR image input and a random sample from the latent distribution, as it is an invertible bijective model.
Note that this model regards as an input or output, rather than as an external conditional prior. Therefore, it is not explicitly conditional on and is fully invertible between HR and LR image pairs. Besides, by approximating the the distribution of with a multi-variate Gaussian distribution, it allows for tractable Jacobian determinant computation, so that the model can be optimized by maximum likelihood estimation (MLE).
3 Model Architecture
The multi-scale architecture proposed in RealNVP is a popular normalizing flow architecture . It consists of levels and at the end of each level, half of the dimensions are factored out. Generally, the factored out dimensions are directly Gaussianized for the computation of negative log-likelihood loss, lacking sufficient modelling of these dimensions. Therefore, based on the multi-scale architecture, we take a further step to model factored out dimensions conditional on the reserved dimensions.
As illustrated in Fig. 2, at each level , is decomposed to low-frequency component and high-frequency component . Then, is modelled by an additional flow that is conditional on the concatenation of tensors from multiple flow levels. By this design, the reconstruction of high-frequency component is hierarchically conditional on frequencies reconstructed from all previous levels. In forward propagation, similar to the depth-first traversal of a binary tree, we first compute , , …, in order. Then, we model the factored out dimensions in a reverse order: , , …, . In inverse propagation, we compute and level by level, from level to level . Note that the determinant of Jacobian of the whole flow can still be efficiently computed, since the conditional relations between and can be represented as an upper triangle block matrix.
The detailed architecture of HCFlow is shown in Fig. 3. For each level, the first layer is the squeeze layer, which transforms the input to a tensor by trading spatial size for number of channels. Then, flow-steps are used for transforming the tensor and decomposing it into different components. More specifically, each flow-step consists of a sequence of three layers: Actnorm layer, invertible convolution layer and affine coupling layer . After that, the split layer is used to evenly split the tensor into two tensors and along the channel dimension. Note that, for the last level, we only keep 3 channels for to make it fit the RGB space of the LR image. Next, is fed to the next level, while is input into an additional flow.
In the -th additional flow, is transformed to the latent variable by flow-steps. Different from above flow-steps, we use conditional affine coupling layer rather than ordinary affine coupling layer to obtain a conditional flow. In particular, we first upscale the conditional feature from level by nearest neighbor interpolation, and concatenate it with . Then, we use a feature extractor to extract image features, which act as the conditional feature for level . Note that the feature extractor only provides scale and shift for an affine coupling during both forward and inverse propagation. Hence the constrains on being invertible and having a tractable Jacobian do not hold for this part. More formally, the hierarchical conditional mechanism of HCFlow is formulated as follows,
where conditional features of different levels are computed in a reverse order, from to .
Particularly, for the last level, we directly model by a Dirac delta function instead of transforming it to another latent variable. This constrains part of the latent space to be the LR image space and implicitly makes the model be conditional on .
4 Training Objectives
When HCFlow is used for image SR, it can be trained by minimizing the negative log-likelihood loss,
which is unsupervised and converges stably. However, in practice, this loss converges slowly and does not provide strong supervision for image SR. To achieve better HR image PSNR, we can add pixel loss on the generated SR image in inverse propagation, leading to a loss function as follows,
where is the ground-truth HR image and is the generated SR image by inputting the ground-truth LR image and sampling the latent variable with temperature . The added pixel loss can help the flow to learn the SR manifold centered around the PSNR-oriented SR image. Furthermore, we can add perceptual loss and GAN loss on the generated SR image to improve the visual quality. This is formulated as,
where is the generated SR image by inputting and sampling with . Note that unlike the pixel loss that uses , is set to 0.8 or 0.9 to preserve the diversity of HR images.
Image rescaling.
Different from image SR, image rescaling aims to recover exactly the same HR image. Following , we regard the invertible HCFlow as an encoder-decoder framework, in which the forward and inverse processes correspond to the encoding and decoding stages, respectively. The loss is as follows,
where is the pixel loss to ensure that, after downscaling and upscaling, the reconstructed image is close to the input . Note that this loss would dramatically decrease the diversity of generated images. Besides, is the pixel loss on the LR image, which guides to be close to the bicubic LR image , so as to generate visually-pleasing LR images in downscaling. The last term is the regularization on the latent variable .
Experiments
We conduct experiments on general image SR, face image SR and image rescaling to show the effectiveness of HCFlow. For image SR experiments, we train the model by three loss combinations: , and . The corresponding learned models are denoted as HCFlow, HCFlow+ and HCFlow++, respectively.
For general image SR (), we set to 2, 13 and 13, respectively. Two 13-block RRDB networks are used as feature extractors. More details on the architecture are provided in the supplementary. The model is trained on the training set of DIV2K and Flickr2K with random flips. The crop patch size and mini-batch size are set to and , respectively. Adam optimizer with and is used for optimization. For HCFLow (with only ), the learning rate is and reduced by half at , , and of iterations. We fine-tune HCFLow+ (with ) for iterations from the pretrained HCFlow. The weight of and are and , respectively. It is worth pointing out that we can achieve even higher PSNR (about 0.2dB) if we train HCFlow+ from scratch. Similarly, we can fine-tune HCFlow++ by further adding and . The loss weighting parameters are , , and .
For face image SR (), are set to 3, 13 and 13, respectively. Three 8-block RRDB networks are used as feature extractors. We train the model on the CelebA training set and test it using first 5,000 images from the testing set. Following , we crop and resize the HR images to the resolution of , and flip them randomly for data augmentation. Other training details are the same as general image SR.
Image rescaling.
For image rescaling (), we set to 2, 8 and 6, respectively. Two 3-block RRDB networks are used as feature extractors. In particular, we use Haar transformation to replace the squeeze layer and remove invertible convolution layers. Details on data preparation and optimizer are the same as general image SR. The learning rate is initialized as and halved at ( iterations in total). The loss weighting parameters are , and , respectively.
Performance evaluation.
Following SRFlow and IRN , we evaluate PSNR and SSIM on the RGB color space for image SR, and on the Y channel of the YCbCr color space for image rescaling. We also use perceptual metric LPIPS and two no-reference metrics, NIQE and BRISQUE , for better visual quality comparison. Pixel standard deviation of 5 samples are used to compare the diversity of results. In addition, Consistency (PSNR between the downscaled SR image and the ground-truth LR image) and LR-PSNR (PSNR between the generated LR image in forward propagation and the ground-truth LR image) are also reported.
2 Ablation Study
To learn a fully invertible flow between HR and LR image pairs, HCFlow constrains part of the latent space to be the LR image space, instead of using the LR image as an external prior. To show the impact, we remove the LR image from the latent space as shown in case 1 and 2 of Table 1. When there is no conditional prior (case 1), the model fails to converge as it does not have enough information for SR . When we replace with ground-truth LR image as a conditional prior (case 2, similar to SRFlow ), it achieves slightly better performance than HCFlow although they have almost the same conditional information. The underlying reason might be that it has a larger latent space than HCFlow.
Ground-truth LR image as a conditional prior.
HCFlow is conditional on and , which are generated during propagation. When we use the ground-truth LR image as a conditional prior to replace (case 3, Table 1), the model achieves similar performance as HCFlow. In fact, since we model the distribution of as a Dirac delta function , would be nearly equal to after model convergence, which is confirmed by the high LR-PSNR. Therefore, conditional on the generated and the external have similar effects.
Hierarchical conditional mechanism.
As shown in case 4 of Table 1, similar to IRN , we assume the LR image and the rest high-frequency component is independent by removing all conditional priors. It yields significantly worse performance because the reconstruction of HR image (high-frequency component) is highly conditional on the LR image (low-frequency component) for image SR. Despite this, it has better results than case 1, as fitting to the LR image space could partly play the role of conditional prior. In case 5, we change from hierarchical conditional mechanism to single-scale conditional mechanism, by removing from level 1. In this case, () is only conditional on from the same level. The performance drops in terms of all kinds of metrics, which shows that the hierarchical conditional mechanism can better model the conditional relations between high-frequency and low-frequency components.
3 Experiments on Image SR
For general image SR (), we compare HCFlow with state-of-the-art CNN-based and flow-based SR models, including the PSNR-oriented EDSR and RRDB , perception-oriented ESRGAN and RankSRGAN , as well as SRFlow . All methods are trained on the same training dataset. From Table 2 and Fig. 4, we have several observations as follows. First, when sampling HR images with temperature , HCFlow acts like a PSNR-oriented model, achieving similar performance as EDSR and RRDB. Adding the HR pixel loss (i.e., HCFlow+) can further improve the PSNR and SSIM by large margins. Second, when , the perceptual metrics of HCFlow are boosted dramatically. With perceptual loss and GAN loss (i.e., HCFlow++), the perceptual metrics are further improved by significant margins in terms of LPIPS and BRISQUE, which is confirmed by the visual results. Note that, unlike ESRGAN and RankSRGAN, the generated HR images of HCFlow++ are still diversified. Third, HCFlow achieves state-of-the-art performance in terms of both quantitative metrics and visual quality. It generates sharp images with few artifacts. In contrast, RRDB and SRFlow tend to produce blurry images, while ESRGAN and RankSRGAN suffer from over-sharpen artifacts and distortions. In addition, HCFlow only has about half of the number of parameters compared with SRFlow.
Face image SR.
We also test HCFlow on face image SR () to show its effectiveness. The compared methods include PSNR-oriented RRDB, perception-oriented ESRGAN and the flow-based SRFlow. As shown in Table 3 and Fig. 5, similar observations as in general image SR can be concluded for face image SR. HCFlow achieves best quantitative and visual performance compared with competing methods. In particular, HCFlow generates sharp faces with natural details, especially on eyes, teeth and hairs. By comparison, other methods suffer from either over-smoothed results or obvious artifacts.
4 Experiments on Image Rescaling
As a unified framework for image SR and image rescaling, HCFlow also achieves state-of-the-art performance in image rescaling. We compare it with three kinds of rescaling methods: (1) bicubic interpolation & state-of-the-art SR models ; (2) encoder-decoder models ; (3) invertible neural networks .
As can be seen from Table 4, when the downscaling process is fixed (i.e., bicubic interpolation), performances of different state-of-the-art SR models are similar and limited. When the downscaling models are optimized for the upscaling models, the results are largely improved. IRN further boosts the performance by joint optimization based on the invertible architecture. Compared with IRN, the proposed HCFlow achieves better performance on all testing datasets with an increased PSNR of . Besides, as shown in Fig. 6, HCFlow can better preserve image details and generates sharper edges than IRN. Since these two models have same number of parameters, HCFlow is more efficient than IRN for image rescaling, which can be attributed to the conditional modelling between high-frequency and low-frequency components.
Conclusion
In this paper, we proposed a unified framework, i.e., hierarchical conditional flow (HCFlow), for both image super-resolution and image rescaling. It learns a fully invertible mapping between HR image and LR image as well as the latent variable. Particularly, we learn the LR image space and design a hierarchical conditional mechanism between the latent variable (high-frequency component) and the LR image (low-frequency component). For image SR, HCFLow is trained by the negative log-likelihood loss, and is further enhanced by pixel loss, perceptual loss and GAN losses for better performance. For image rescaling, it is trained as an encoder-decoder framework, where the forward and inverse progresses are jointly optimized. Experiments demonstrate that HCFlow achieves state-of-the-art performance on general image SR, face image SR and image rescaling, in terms of both quantitative metrics and visual quality.
Acknowledgements We thank Dr. Suryansh Kumar for helpful discussion. This work was partially supported by the ETH Zurich Fund (OK), a Huawei Technologies Oy (Finland) project, the China Scholarship Council and a Microsoft Azure grant. Special thanks goes to Yijue Chen.