The 2018 PIRM Challenge on Perceptual Image Super-resolution

Yochai Blau, Roey Mechrez, Radu Timofte, Tomer Michaeli, Lihi Zelnik-Manor

Introduction

The past few years have seen a major performance leap in single-image super-resolution (SR), both in terms of reconstruction accuracy (as measured e.g., by PSNR, SSIM) and in terms of visual quality (as rated by human observers) . However, the more SR methods advanced, the more it has become evident that reconstruction accuracy and perceptual quality are typically in disagreement with each other. That is, models which excel at minimizing the reconstruction error tend to produce visually unpleasing results, while models that produce results with superior visual quality are rated poorly by distortion measures like PSNR, SSIM, IFC, etc. (see Fig. 1). Recently, it has been shown that this disagreement cannot be completely resolved by seeking for better distortion measures . Namely, there is a fundamental tradeoff between the ability to achieve low distortion and low deviation from natural image statistics, no matter what full-reference dissimilarity criterion is used to measure distortion.

These observations caused the formation of two distinct research trends (see Fig. 2). The first is aimed at improving the reconstruction accuracy according to popular full-reference distortion metrics, and the second targets high perceptual quality. While reconstruction accuracy can be precisely quantified, perceptual quality is often estimated through user studies, in which, due to practical limitations, each user is typically exposed to only a small number of methods and/or a small number of images per method. Therefore, reports on perceptual quality are often inaccurate and hard to reproduce. As a result, novel methods cannot be easily compared to their predecessors in terms of perceptual quality, and existing benchmarks and challenges (e.g., NTIRE ) focus mostly on quantifying reconstruction accuracy, using e.g., PSNR/SSIM. As perceptually-aware super-resolution is gaining increasing attention in recent years, there is a need for a benchmark for evaluating perceptual-quality driven algorithms.

The 2018 PIRM challenge on perceptual super-resolution took part in conjunction with the 2018 Perceptual Image Restoration and Manipulation (PIRM) workshop. This challenge compared and ranked perceptual super-resolution algorithms. In contrast to previous challenges, the evaluation was performed in a perceptual-quality aware manner, as suggested in . Specifically, we define perceptual quality as the visual quality of the reconstructed image regardless of its similarity to any ground-truth image. Namely, it is the extent to which the reconstruction looks like a valid natural image. Therefore, we measured the perceptual quality of the reconstructed images using no-reference image quality measures, which do not rely on the ground-truth image.

Although the main motivation of the challenge is to promote algorithms that produce images with good perceptual quality, similarity to the ground truth images is obviously also of importance. For example, perfect perceptual quality can be achieved by randomly drawing natural images that have nothing to do with the input images. Such a scheme would score quite poorly in terms of reconstruction accuracy. We therefore evaluate algorithms on a 2-dimensional plane, where one axis is the full-reference root mean squared error (RMSE) distortion, and the second axis is a perceptual index which combines the no-reference image quality measures of and . This approach jointly quantifies accuracy and perceptual quality, thus enabling perceptual-driven methods to compete alongside algorithms that target PSNR maximization. PIRM is therefore the first established benchmark for perceptual-quality driven image restoration, which will hopefully be extended to other perceptual computer-vision tasks in the future.

The outcomes arising from this challenge are manifold:

∙\bullet Participants introduced algorithms which well-improve upon the state of the art in perceptual SR. The submitted methods incorporated novelties in optimization objectives (losses), conv-net architectures, generative adversarial net (GAN) variants, training schemes and more. These enabled to impressively surpass the performance of baselines, such as EnhanceNet and CX . The results are presented in Section 4, and the main novelties are discussed in Section 6.

∙\bullet We validate our chosen perceptual index through a human-opinion study, and find that it is highly correlated with the ratings of human observers. This provides empirical evidence that no-reference image quality measures can faithfully assess perceptual quality. The results of the human-opinion study are presented in Section 4.1.

∙\bullet We also test the agreement of many other commonly used image quality measures with the human-opinion scores, and find that most of them are either uncorrelated or anti-correlated. This shows that most existing schemes for evaluating image restoration algorithms cannot be used to quantify perceptual quality. The results of this analysis are presented in Section 5.

∙\bullet The challenge results provide insights on the trade-off between perception and distortion (suggested and analyzed in ). In particular, at the low-distortion regime, participants showed considerable improvements in perceptual quality over methods that excel in RMSE (e.g. EDSR ), while sacrificing only a small increase in RMSE. This indicates that the tradeoff is severe in this regime. Furthermore, at the good perceptual quality regime, participants were able to improve both in perceptual quality and in distortion, over state-of-the-art perceptual SR methods (e.g.E-Net ). This indicates that previous methods were quite far from the theoretical perception-distortion bound discussed in .

Perceptual Super Resolution

These perceptual SR methods have established a fresh research direction which is producing algorithms with superior perceptual quality. However, in all works, this has come at the cost of a substantial decrease in PSNR and SSIM values, indicating that these common distortion measures do not faithfully quantify the perceptual quality of SR methods . As such, perceptual SR algorithms cannot participate in any challenge or benchmark based on these standard measures (e.g., NTIRE ), and cannot be compared or ranked using these common metrics.

The PIRM Challenge on Perceptual SR

The PIRM challenge is the first to compare and rank perceptual image super-resolution. The essential difference compared to previous challenges is the novel evaluation scheme which is not based solely on common distortion measures such as PSNR/SSIM.

The challenge task is 4×4\times super-resolution of a single image which was down-sampled with a bicubic kernel.

Validation and testing of the submitted methods were performed on two sets of 100 images eachThe validation set was used throughout the challenge for model development, and the test set was released a week before the challenge ended for assessing the final results.. These images cover diverse contents, including people, objects, environments, flora, natural scenery, etc. Participants did not have access to the high-res ground truth images during the challenge, and these images were not available on any online source prior to the challenge. These image sets (high and low resolution) are now available onlinehttps://pirm.github.io. Datasets for model training were chosen by the participants.

The evaluation scheme is based on , which proposed to evaluate image restoration algorithms on the perception-distortion plane (see Fig. 3). The rationale of this method is shortly explained in the Introduction.

In the PIRM challenge, the perception-distortion plane was divided into three regions by setting thresholds on the RMSE values (regions 1/2/31/2/3 were defined by RMSE≤11.5/12.5/16\text{RMSE}\leq 11.5/12.5/16 respectively, see Fig. 3). In each region, the goal was to obtain the best mean perceptual quality. That is, participants attempted to move as downwards as possible in the perception-distortion plane. The perception index (PI) we chose for the vertical axis combines the no-reference image quality measures of Ma et al. and NIQE as

Notice that in this setting, a lower perceptual index indicates better perceptual quality. The RMSE was computed as the square-root of the mean-squared-error (MSE) of all pixels in all imagesNote that this is not the mean of the RMSEs of the images, but rather the square-root of the images’ mean MSE., that is

where xiHRx_{i}^{\text{HR}} and xiESTx_{i}^{\text{EST}} are the iith ground truth and estimated images respectively, NiN_{i} is the number of pixels in xiHRx_{i}^{\text{HR}}, and MM is the number of images in the test set. Both the RMSE and the PI were computed on the y-channel after removing a 44-pixel border. We encouraged participants to submit methods for all three regions, and indeed many did (see Table 1).

Challenge Results

Twenty-one teams participated in the test phase of the challenge. Table 1 reports the top scoring teams in each region, where the team members and affiliations can be found in Appendix 0.A. Figure 4(a) plots all test phase submissions on the perception-distortion plane (teams were allowed up to 10 final submissions). Figure 4(b) shows the correlation between our perceptual index (PI) and human-opinion-scores on the top 10 submissions (see details in Sec. 5). The high correlation justifies our choice of definition of the PI. In Fig. 5 we compare the visual outputs of several top methods in each region (the number in the method’s name indicates the region of the submission), where additional visual comparisons can be found in Appendix 0.C. A table with the scores of all participating teams in each region can be found in Appendix 0.B.

The submitted algorithms exceed the performance of previous SR methods in all regions, pushing forward the state-of-the-art in perceptual SR. In Region 33, challenge submissions outperform the EnhanceNet baseline, as well as the recently proposed CX algorithm. Notice that several submissions improve upon the baselines in both perceptual quality and reconstruction accuracy, which are both important. In Region 22, the top submissions present fairly good perceptual quality with a far lower distortion than the methods in Region 33. Such methods could prove advantageous in applications where reconstruction accuracy is valuable. Inspection of the Region 11 results reveals that participants obtained a significant improvement in the PI (45%45\%) w.r.t. the EDSR baseline with only a small increase in the RMSE (7%,0.777\%,0.77 gray-levels per-pixel).

The results provide insights on the tradeoff between perceptual quality and distortion, which is clearly noticed when progressing from Region 11 to Region 33. First, the tradeoff appears to be stronger in the low distortion regime (Region 11), implying that PSNR maximization can have damaging effects in terms of perceptual quality. In the high perceptual quality regime (Region 33), notice that beyond some point, increasing the RMSE allows only slight improvement in the perceptual quality. This indicates that it is possible to achieve perceptual quality similar to that of the current state-of-the-art methods with considerably lower RMSE values.

We validate the challenge results with a human-opinion study. Thirty-five raters were each shown the outputs of 12 algorithms (10 top challenge submissions, 2 baselines) on 20 images (240 images per rater). For each image, they were asked to rate how realistic the image looked on a scale of 1−41-4 which corresponds to: 11-Definitely fake, 22-Probably fake, 33-Probably real, and 44-Definitely real. We made it clear that “real” corresponds to a natural image and “fake” corresponds to the output of an algorithm. This scale tests how natural the outputs look. Note that users were not exposed to the original “ground truth” images, therefore this study does not test distortion in any way, but rather only perceptual quality. The mean human-opinion-scores are shown in Fig. 6.

The human-opinion study validates that the challenge submissions surpassed the performance of state-of-the-art baselines by significant margins. Region 33 submissions, and even Region 22 submissions, are considered notably better than EnhanceNet by human raters. Region 11 submissions were rated far better in visual quality compared to EDSR (with only a slight increase in RMSE). The tradeoff between perceptual quality and distortion is once more revealed, as the best attainable perceptual quality increases with the increase in RMSE. Note that while the PI is well correlated with the human-opinion-scores on a coarse scale (in between regions), it is not always well-correlated with these scores on a finer scale (rankings within the regions), which can be seen when comparing the rankings in Table 1 and Fig. 6. This highlights the urgent need for better perceptual quality metrics, a point which is further analyzed in Section 5.

Figure 7 shows the normalized histogram of votes per method. Notice that all methods fail to achieve a large percentage of “definitely real” votes, indicating that there is still much to be done in perceptual super-resolution. In all submitted results, there tend to appear unnatural features in the reconstructions (at 4×4\times magnification), which degrade the perceptual quality. Notice that the outputs of EDSR, a state-of-the-art algorithm in terms of distortion, are mostly voted as “definitely fake”. This is due to the aggressive averaging causing blurriness as a consequence of optimizing for distortion.

2 Not all images are created equal

The results presented in the previous sections show the general trends when averaging over a set of images. Interestingly, when examining single images, there can be quite a variability in SR results. First, there are images which are much easier to super-resolve than others. In such a scenario, the outputs of all SR methods tend towards high perceptual quality. Such an example can be seen on the left side of Fig. 8, where the outputs of all methods on the “grafity” image are rated fairly higher compared to the “mountain” image. In both it seems advantageous to move towards region 33, but the SR of texture-less images (such as “grafity”) will generally produce visually pleasing results. Another variation from the average trend are images which include more structure than texture. On such images, it seems that methods from region 11 which prefer accuracy succeed in maintaining large-scale structures, as opposed to generative-based methods from region 33 which tend to distort structures and often produce visually unpleasing results. For example, on the “building” image on the right side of Fig. 8, the outputs of EDSR are visually pleasing while the outputs of region 33 methods are rated unsatisfactory. However, for images with fine unstructured details such as the “carved stone” image, it is beneficial to move towards region 33. This calls for novel methods, which can either adaptively favor structure preservation vs. texture reconstruction, or employ generative models capable of outputing large-scale structured regions.

Analyzing Quality Measures

The lack of a faithful criterion for assessing the perceptual quality of images is restricting progress in perceptually-aware image reconstruction and manipulation tasks. The current main tool for comparing methods are human-opinion studies, which are hardly reproducible, making it practically impossible to systematically compare methods and assess progress. Here, we analyze the relation between existing image quality metrics and human-opinion scores, concluding which metrics are best for quantifying perceptual quality. In Fig. 9, we plot the mean-opinion scores of the methods included in the human-opinion study vs. the mean score according to the common full-reference measures RMSE, SSIM , IFC , and LPIPS , as well as the no-reference methods by Ma et al. , NIQE , BRISQUE and the PI defined by (1). For each measure, we report Spearman’s correlation coefficient with the raters’ mean opinion scores, and also plot the corresponding least-squares linear fit.

As seen in Fig. 9, RMSE, SSIM and IFC, which are widely used for evaluating the quality of image reconstruction algorithms, are anti-correlated with perceptual quality and thus inappropriate for evaluating it. Ma et al. and BRISQUE show moderate correlation with human-opinion-scores, while LPIPS, NIQE and PI are highly correlated, with PI being the most correlated.

The bottom pane of Fig. 9 focuses on the high-perceptual quality regime, where it is important to distinguish between methods and correctly rank them. Metrics which excel in this regime will allow to assess progress in perceptual SR and to systematically compare methods. This is done by zooming in on the region of mean-opinion-score above 2.32.3 (a new least-squares linear fit appears in magenta). These plots reveal that LPIPS, Ma et al. and BRISQUE fail to faithfully quantify the perceptual quality in this regime. The only methods capable of correctly evaluating the perceptual quality of perceptually-aware SR algorithms are NIQE and PI (which is a combination of NIQE and Ma). Note that we also tested the full-reference measures VIF , FSIM and MS-SSIM , and the no-reference measures CORNIA and BLIINDS , which all failed to correctly assess the perceptual qualityVIF, FSIM, MS-SSIM and CORNIA were anti-correlated with the mean-opinion-scores. BLIINDS was moderately correlated, but failed in the high perceptual quality regime (similar to BRISQUE)..

We also analyze the correlation between human-opinion scores and common image quality measures on a single image. In Fig. 10 we plot the scores for outputs of each tested challenge method on all 4040 tested images (480480 images altogether), where we average only over different human raters. To eliminate the variations between images (see Section 4.2), we first subtract the mean score of each image (over different raters) for both the human-opinion scores and the image quality measures. As can be seen, theses results are similar in trend to the results presented in Fig. 9.

Current Trends in Perceptual Super Resolution

All twenty-one groups who participated in the PIRM SR challenge, submitted algorithms based on deep nets. We next shortly review the current trends reflected in the submitted algorithms, in terms of three main aspects: the loss functions, the architectures, and methods to traverse the perception-distortion tradeoff. Note that the scope of this paper is not to review the field of SR, but rather to summarize the leading trends in the PIRM SR challenge. Additional details on the submitted methods can be found in the PIRM workshop proceedings.

A different approach that achieved high perceptual quality is transferring texture by training with the Gram loss , and without adversarial training. These participants show that standard texture transfer can be further improved by controlling the process using homogeneous semantic regions.

Submissions also applied other distortion functions, including the MS-SSIM loss function to emphasize a more structural distortion goal, Discrete Cosine Transform (DCT) based loss function and L1 norm between image gradients which were suggested in order overcome the smoothing effect of the MSE loss.

2 Architecture

The second crucial component of submissions is the network architecture. Overall, most participating teams adopted state-of-the-art architectures from successful PSNR-maximization based SR methods and replaced the loss function. The main trend is to use the EDSR network architecture for the generator and the SRGAN architecture for the discriminator. Wang et al. suggested to replace the residual block of EDSR with the Residual-in-Residual Dense Block (RRDB), which combines multi-level residual networks and dense connections. RRDR enables the use of deeper models, and as a result, improves the recovered textures. Others used Deep Back-Projection Networks (DBPN) , Enhanced Upscale Modules (EUSR) , and Multi-Grid-Back-Projection (MGBP) .

3 Traversing the perception-distortion tradeoff

The tradeoff between perceptual quality and distortion raises the question of how to control the compromise between these two objectives. The importance of this question is two-fold: first, the optimal working point along the perception-distortion curve is domain specific and moreover it is image specific. Second, it is hard to predict the final working point, especially when the full objective is complex and when adversarial training is incorporated. Below we elaborate on four possible solutions (see pros and cons in Table 2):

Retrain the network for each working point. This can be done by modifying the magnitude of the loss terms (e.g. adversarial and distortion losses).

Interpolate between output images of two pretrained networks (in the pixel domain). For example, by using soft thresholding .

Interpolate between the parameters of two networks with the same architecture but different loss. This allows to generate a third network that is easy to control (see for details).

Control the tradeoff with an additional network input. For example, added noise to the input in order to traverse along the curve by changing the noise level at test time.

Conclusions

The 2018 PIRM challenge is the first benchmark for perceptual-quality driven SR algorithms. The novel evaluation methodology used in this challenge enabled the assessment and ranking of perceptual SR methods along-side with those which target PSNR maximization. With this evaluation scheme, we compared the submitted algorithms with existing baselines, which revealed that the proposed methods push forward this field’s state-of-the-art. A thorough study of the capability of common image quality measures to capture the perceptual quality of images was conducted. This study exposed that most common image quality measures are inadequate of quantifying perceptual quality.

We conclude this report by pointing to several challenges in the field of perceptual SR, which should be the focus of future work. While we have witnessed major improvements over the past several years, in challenging scenarios such as 44x SR, the outputs of current methods are generally unrealistic to human observers. This highlights that there is still much to be done to achieve high-quality perceptual SR images. Most common image quality measures fail to quantify the perceptual quality of SR methods, and there is still much room for improvement in this essential task. Perceptual-quality driven algorithms have yet to appear for the real-world scenario of blind SR. The perceptual quality objective, which has gained much attention for the SR task, should also gain attention for other image restoration tasks e.g. deblurring. Finally, since a tradeoff between reconstruction accuracy and perceptual quality exists, schemes for controlling the compromise between the two can lead to adaptive SR schemes. This may promote new ways of quantifying the performance of SR algorithms, for instance, by measuring the area-under-the-curve in the perception-distortion plane.

The 2018 PIRM Challenge on Perceptual SR was sponsored by Huawei and Mediatek.

References

Appendix 0.A Participating teams

Appendix 0.B Test phase results

Appendix 0.C More results