The Perception-Distortion Tradeoff
Yochai Blau, Tomer Michaeli
Introduction
The last decades have seen continuous progress in image restoration algorithms (e.g. for denoising, deblurring, super-resolution) both in visual quality and in distortion measures like peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) . However, in recent years, it seems that the improvement in reconstruction accuracy is not always accompanied by an improvement in visual quality. In fact, and perhaps counter-intuitively, algorithms that are superior in terms of perceptual quality, are often inferior in terms of e.g. PSNR and SSIM . This phenomenon is commonly interpreted as a shortcoming of the existing distortion measures , which fuels a constant search for alternative “more perceptual” criteria.
In this paper, we offer a complementary explanation for the apparent tradeoff between perceptual quality and distortion measures. Specifically, we prove that there exists a region in the perception-distortion plane, which cannot be attained regardless of the algorithmic scheme (see Fig. 1). Furthermore, the boundary of this region is monotone. Therefore, in its proximity, it is only possible to improve either perceptual quality or distortion, one at the expense of the other. The perception-distortion tradeoff exists for all distortion measures, and is not only a problem of the mean-square error (MSE) or SSIM criteria.
Let us clarify the difference between distortion and perceptual quality. The goal in image restoration is to estimate an image from its degraded version (e.g. noisy, blurry, etc.). Distortion refers to the dissimilarity between the reconstructed image and the original image . Perceptual quality, on the other hand, refers only to the visual quality of , regardless of its similarity to . Namely, it is the extent to which looks like a valid natural image. An increasingly popular way of measuring perceptual quality is by using real-vs.-fake user studies, which examine the ability of human observers to tell whether is real or the output of an algorithm (similarly to the idea underlying generative adversarial nets ). Therefore, perceptual quality can be defined as the best possible probability of success in such discrimination experiments, which as we show, is proportional to the distance between the distribution of reconstructed images and that of natural images.
Based on these definitions of perception and distortion, we follow the logic of rate-distortion theory . That is, we seek to characterize the behavior of the best attainable perceptual quality (minimal deviation from natural image statistics) as a function of the maximal allowable average distortion, for any estimator. This perception-distortion function (wide curve in Fig. 1) separates between the attainable and unattainable regions in the perception-distortion plane and thus describes the fundamental tradeoff between perception and distortion. Our analysis shows that algorithms cannot be simultaneously very accurate and produce images that fool observers to believe they are real, no matter what measure is used to quantify accuracy. This tradeoff implies that optimizing distortion measures can be not only ineffective, but also potentially damaging in terms of visual quality. This has been empirically observed e.g. in , but was never established theoretically.
From the standpoint of algorithm design, we show that generative adversarial nets (GANs) provide a principled way to approach the perception-distortion bound. This gives theoretical support to the growing empirical evidence of the advantages of GANs in image restoration .
The perception-distortion tradeoff has major implications on low-level vision. In certain applications, reconstruction accuracy is of key importance (e.g. medical imaging). In others, perceptual quality may be preferred. The impossibility of simultaneously achieving both goals calls for a new way for evaluating algorithms: By placing them on the perception-distortion plane. We use this new methodology to conduct an extensive comparison between recent super-resolution (SR) methods, revealing which SR methods lie closest to the perception-distortion bound.
Distortion and perceptual quality
Distortion and perceptual quality have been studied in many different contexts, and are sometimes referred to by different names. Let us briefly put past works in our context.
Given a distorted image and a ground-truth reference image , full-reference distortion measures quantify the quality of by its discrepancy to . These measures are often called full reference image quality criteria because of the reasoning that if is similar to and is of high quality, then is also of high quality. However, as we show in this paper, this logic is not always correct. We thus prefer to call these measures distortion or dissimilarity criteria.
2 Perceptual quality
The perceptual quality of an image is the degree to which it looks like a natural image, and has nothing to do with its similarity to any reference image. In many image processing domains, perceptual quality has been associated with deviations from natural image statistics.
Perceptual quality is commonly evaluated empirically by the mean opinion score of human subjects . Recently, it has become increasingly popular to perform such studies through real vs. fake questionnaires . These test the ability of a human observer to distinguish whether an image is real or the output of some algorithm. The probability of success of the optimal decision rule in this hypothesis testing task is known to be (see Appendix A in the Supplementary Material)
where is the total-variation (TV) distance between the distribution of images produced by the algorithm in question, and the distribution of natural images . Note that decreases as the deviation between and decreases, becoming (no better than a coin toss) when .
No-reference quality measures
Perceptual quality can also be measured by an algorithm. In particular, no-reference measures quantify the perceptual quality of an image without depending on a reference image. These measures are commonly based on estimating deviations from natural image statistics. The works proposed perceptual quality indices based on the KL divergence between the distribution of the wavelet coefficients of and that of natural scenes. This idea was further extended by the popular methods DIIVINE , BRISQUE , BLIINDS-II and NIQE , which quantify perceptual quality by various measures of deviation from natural image statistics in the spatial, wavelet and DCT domains.
GAN-based image restoration
Most recently, GAN-based methods have demonstrated unprecedented perceptual quality in super-resolution , inpainting , compression , deblurring and image-to-image translation . This was accomplished by utilizing an adversarial loss, which minimizes some distance between the distribution of images produced by the generator and the distribution of images in the training dataset. A large variety of GAN schemes have been proposed, which minimize different distances between distributions. These include the Jensen-Shannon divergence , the Wasserstein distance , and any -divergence .
Single image quality vs. image ensemble quality
Here is the Kullback-Leibler divergence and denotes entropy. The second term in this decomposition discourages diversity. Choosing to drop it, results in the distributional-divergence based quality measures described above.
Problem formulation
In statistical terms, a natural image can be thought of as a realization from the distribution of natural images . In image restoration, we observe a degraded version relating to via some conditional distribution . In this paper we focus on non-invertible settingsBy invertible we mean that the support of is a singleton for almost all ’s (see Appendix C for a formal definition)., where cannot be estimated from with zero error. This is typically the case in denoising, deblurring, inpaitning, super-resolution, etc. Given , an image restoration algorithm produces an estimate according to some distribution . Note that this description is quite general in that it does not restrict the estimator to be a deterministic function of . This problem setting is illustrated in Fig. 2.
Given a full-reference dissimilarity criterion , the average distortion of an estimator is given by
where the expectation is over the joint distribution . This definition aligns with the common practice of evaluating average performance over a database of degraded natural images. We assume that the dissimilarity criterion is such that with equality when . Note that some distortion measures, e.g. SSIM, are actually similarity measures (higher is better), yet can always be inverted (and shifted) to become dissimilarity measures.
As discussed in Sec. 2.2, the perceptual quality of an estimator (as quantified e.g. by real vs. fake human opinion studies) is directly related to the distance between the distribution of its reconstructed images, , and the distribution of natural images, . We thus define the perceptual quality index (lower is better) of an estimator as
where is some divergence between distributions that satisfies with equality if , e.g. the KL divergence, TV distance, Wasserstein distance, etc. It should be pointed out that the divergence function which best relates to human perception is a subject of ongoing research. Yet, our results below hold for (nearly) any divergence.
Notice that the best possible perceptual quality is obtained when the outputs of the algorithm follow the distribution of natural images (i.e. ). In this situation, by looking at the reconstructed images, it is impossible to tell that they were generated by an algorithm. However, not every estimator with this property is necessarily accurate. Indeed, we could achieve perfect perceptual quality by randomly drawing natural images that have nothing to do with the original ground-truth images. In this case the distortion would be quite large.
Our goal is to characterize the tradeoff between (3) and (4). But let us first see why minimizing the average distortion (3), does not necessarily lead to a low perceptual quality index (4). We start by illustrating this with the square-error distortion and the distortion (where is Kronecker’s delta). We specifically illustrate that those measures are generally not distribution preserving in the following sense.
We say that a distortion measure is distribution preserving at if the estimator minimizing the mean distortion (3) satisfies .
More details about those examples are provided in Appendix B (Supplementary Material). We then proceed to discuss this phenomenon for arbitrary distortions in Sec. 3.3.
and is independent of (see Fig. 3). In this setting, the MMSE estimate is given by
Notice that can take any value in the range , whereas can only take the discrete values . Thus, clearly, is very different from , as illustrated in Fig. 3. This demonstrates that minimizing the MSE distortion does not generally lead to .
The same intuition holds for images. The MMSE estimate is an average over all possible explanations to the measured data, weighted by their likelihoods. However the average of valid images is not necessarily a valid image, so that the MMSE estimate frequently “falls off” the natural image manifold . This leads to unnatural blurry reconstructions, as illustrated in Fig. 4. In this experiment, is a image comprising smaller digit images. Each digit is chosen uniformly at random from a dataset comprising K images from the MNIST dataset and an additional K blank images. The degraded image is a noisy version of . As can be seen, the MMSE estimator produces blurry reconstructions, which do not follow the statistics of the (binary) images in the dataset.
2 The 0−1010-1 distortion
The discussion above may give the impression that unnatural estimates are mainly a problem of the square-error distortion, which causes averaging. One way to avoid averaging, is to minimize the binary loss, which restricts the estimator to choose only from the set of values that can take. In fact, the minimum mean distortion is attained by the maximum-a-posteriori (MAP) rule, which is very popular in image restoration. However, as we exemplify next, the distribution of the MAP estimator also deviates from . This behavior has also been studied in .
Consider again the setting of (5). In this case, the MAP estimate is given by
where is as in (7). Now, it can be easily verified that when , we have . Namely, the MAP estimator never predicts the value . Therefore, in this case, the distribution of the estimate is
which is obviously different from of (5) (see Fig. 3).
This effect can also be seen in the experiment of Fig. 4. Here, the MAP estimates become increasingly dominated by blank images as the noise level rises, and thus clearly deviates from the underlying prior distribution.
3 Arbitrary distortion measures
We saw that neither the square-error nor the loss are distribution preserving. That is, their minimization does not generally lead to (i.e. perfect perceptual quality). However these two examples do not yet preclude the existence of a distribution preserving distortion measure. Does there exist a measure whose minimization is guaranteed to lead to ? If we limit ourselves to one single setting, then the answer may be positive. For example, in the setting of Fig. 3, if of (5) equals , then the loss is distribution preserving as its minimization leads to an estimate satisfying . This illustrates that a distortion measure may be distribution preserving for certain underlying distributions but not for others.
However, from a practical standpoint, we typically want our distortion measure to be adequate in more than one single setting. For example, if our goal is to train a neural network to perform denoising, then it is reasonable to expect that the same distortion measure be equally adequate as a loss function for different noise levels. In fact, we may also want to use the same distortion measure across different tasks (e.g. super-resolution, deblurring, inpainting). The interesting question is, therefore, whether there exists a stably distribution preserving distortion measure.
As we show next, if the degradation is non-invertible, then no distortion metric can be stably distribution preserving (see proof in Appendix C).
If defines a non-invertible degradation, then is not a stably distribution preserving distortion at .
The perception-distortion tradeoff
We saw that for any distortion measure, a low distortion does not generally imply good perceptual-quality. An interesting question, then, is: What is the best perceptual quality that can be attained by an estimator with a prescribed distortion level?
The perception-distortion function of a signal restoration task is given by
where is a distortion measure and is a divergence between distributions.
In words, is the minimal deviation between the distributions and that can be attained by an estimator with distortion . To gain intuition into the typical behavior of this function, consider the following example.
Suppose that , where and are independent. Take to be the square-error distortion and to be the KL divergence. For simplicity, let us focus on estimators of the form . In this case, we can derive a closed form solution to Eq. (10) (see Appendix D), which is plotted for several noise levels in Fig. 5. As can be seen, the minimal attainable drops as the maximal allowable distortion (MSE) increases. Furthermore, the tradeoff is convex and becomes more severe at higher noise levels .
In general settings, it is impossible to solve (10) analytically. However, it turns out that the behavior seen in Fig. 5 is typical, as we show next (see proof in Appendix E).
Assume the problem setting of Section 3. If of (4) is convex in its second argumentThat is, for any three distributions and any ., then the perception-distortion function of (10) is
Note that Theorem 2 requires no assumptions on the distortion measure . This implies that a tradeoff between perceptual quality and distortion exists for any distortion measure, including e.g. MSE, SSIM, square error between VGG features , etc. Yet, this does not imply that all distortion measures have the same perception-distortion function. Indeed, as we demonstrate in Sec. 6, the tradeoff tends to be less severe for distortion measures that capture semantic similarities between images.
The convexity of implies that the tradeoff is more severe at the low-distortion and at the high-perceptual-quality extremes. This is particularly important when considering the TV divergence which is associated with the ability to distinguish between real vs. fake images (see Sec. 2.2). Since is steeper at the low-distortion regime, any small improvement in distortion for an algorithm whose distortion is already low, must be accompanied by a large degradation in the ability to fool a discriminator. Similarly, any small improvement in the perceptual quality of an algorithm whose perceptual index is already low, must be accompanied by a large increase in distortion. Let us comment that the assumption that is convex, is not very limiting. For instance, any -divergence (e.g. KL, TV, Hellinger, ) as well as the Renyi divergence, satisfy this assumption . In any case, the function is monotonically non-increasing even without this assumption.
Several past works attempted to answer the question: What is the minimal attainable distortion in various restoration tasks? . This corresponds to the value
which is the horizontal coordinate of the leftmost point on the perception-distortion function. However, as the minimum distortion estimator is generally not distribution preserving (Sec. 3.3), an important complementary question is: What is the minimal distortion that can be attained by an estimator having perfect perceptual quality? This corresponds to the value
which is the horizontal coordinate of the point where the perception-distortion function first touches the horizontal axis (see Fig. 6).
Observe that perfect perceptual quality () is always attainable, for example by drawing from independently of the input . This method, however, ignores the input and is thus not good in terms of distortion. It turns out that perfect perceptual quality can generally be achieved with a significantly lower MSE distortion, as we show next (see proof in Appendix F).
For the square error distortion ,
where and are defined by (11) and (12), respectively. This bound is attained by the estimator defined through
which achieves and has an MSE of .
In simple words, Theorem 3 states that one would never need to sacrifice more than dB in PSNR to obtain perfect perceptual quality. This can be achieved by drawing from the posterior distribution . Interestingly, such a degradation was indeed incurred by all super-resolution methods that achieved state-of-the-art perceptual quality to date. This can be seen in Fig. 9, where the RMSE of the algorithms with the lowest perceptual index is nearly a factor of larger than the RMSE of the methods with the lowest RMSE (see also ). However, note that this bound is generally not tight. For example, in the scalar Gaussian toy example of Fig. 5, can be quite smaller than , depending on the noise level.
2 Connection to rate-distortion theory
The perception-distortion tradeoff is closely related to the well-established rate-distortion theory . This theory characterizes the tradeoff between the bit-rate required to communicate a signal, and the distortion incurred in the signal’s reconstruction at the receiver. More formally, the rate-distortion function of a signal is defined by
where is the mutual information between and .
There are, however, several key differences between the two tradeoffs. First, in rate-distortion the optimization is over all conditional distributions , i.e. given the original signal. In the perception-distortion case, the estimator has access only to the degraded signal , so that the optimization is over the conditional distributions , which is more restrictive. In other words, the perception-distortion tradeoff depends on the degradation , and not only on the signal’s distribution (see Example 1). Second, in rate-distortion the rate is quantified by the mutual information , which depends on the joint distribution . In our case, perception is quantified by the similarity between and , which does not depend on their joint distribution. Lastly, mutual information is inherently convex, while the convexity of the perception-distortion curve is guaranteed only when is convex.
While the two tradeoffs are different, it is important to note that perceptual quality does play a role in lossy compression, as evident from the success of recent GAN based compression schemes . Theoretically, its effect can be studied through the rate-distortion-perception function , which is an extension of the rate-distortion function (15) and the perception-distortion function (10), characterizing the triple tradeoff between rate, distortion, and perceptual quality.
Traversing the tradeoff with a GAN
There exists a systematic way to design estimators that approach the perception-distortion curve: Using GANs. Specifically, motivated by , restoration problems can be approached by modifying the loss of the generator of a GAN to be
Besides the MMSE estimator, Figure 7 also includes the MAP estimator, the random draw estimator (which ignores the noisy image ), and the conditional draw estimator of (14). The perceptual quality of these estimators is evaluated, as above, by the final loss of the WGAN discriminator , trained (without a generator) to distinguish between the estimators’ outputs and images from the dataset. Note that the denoising WGAN estimator (D) achieves the same distortion as the MAP estimator, but with far better perceptual quality. Furthermore, it achieves nearly the same perceptual quality as the random draw estimator, but with a significantly lower distortion.
Practical method for evaluating algorithms
Certain applications may require low-distortion (e.g. in medical imaging), while others may prefer superior perceptual quality. How should image restoration algorithms be evaluated, then?
We say that Algorithm A dominates Algorithm B if it has better perceptual quality and less distortion.
Note that if Algorithm A is better than B in only one of the two criteria, then neither dominates nor dominates . Therefore, among a group of algorithms, there may be a large subset which can be considered equally good.
We say that an algorithm is admissible among a group of algorithms, if it is not dominated by any other algorithm in the group.
As shown in Figure 8, these definitions have very simple interpretations when plotting algorithms on the perception-distortion plane. In particular, the admissible algorithms in the group, are those which lie closest to the perception-distortion bound.
As discussed in Sec. 2, distortion is measured by full-reference (FR) metrics, e.g. . The choice of the FR metric, depends on the type of similarities we want to measure (per-pixel, semantic, etc.). Perceptual quality, on the other hand, is ideally quantified by collecting human opinion scores, which is time consuming and costly . Instead, the divergence can be computed, for instance by training a discriminator net (see Sec. 5). However, this requires many training images and is thus also time consuming. A practical alternative is to utilize no-reference (NR) metrics, e.g. , which quantify the perceptual quality of an image without a corresponding original image. In scenarios where NR metrics are highly correlated with human mean-opinion-scores (e.g. super-resolution ), they can be used as a fast and simple method for approximating the perceptual quality of an algorithmIn scenarios where NR metrics are inaccurate (e.g. blind deblurring with large blurs ), the perceptual metric should be human-opinion-scores or the loss of a discriminator trained to distinguish the algorithms’ outputs from natural images..
We use this approach to evaluate SR algorithms in a magnification task, by plotting them on the perception-distortion plane (Fig. 9). We measure perceptual quality using the NR metric NIQE , which was shown to correlate well with human opinion scores in a recent SR challenge (see Appendix H for experiments with the NR metrics BRISQUE , BLIINDS-II and the recent NR metric by Ma et al. ). We measure distortion by the five common FR metrics RMSE, SSIM , MS-SSIM , IFC and VIF , and additionally by the recent metric (the distance in the feature space of a VGG net) . To conform to previous evaluations, we compute all metrics on the y-channel after discarding a 4-pixel border (except for VGG2,2, which is computed on RGB images). Comparisons on color images can be found in Appendix H. The algorithms are evaluated on the BSD100 dataset . The evaluated algorithms include: A+ , SRCNN , SelfEx , VDSR , Johnson et al. , LapSRN , Bae et al. (“primary” variant), EDSR , SRResNet variants which optimize MSE and , SRGAN variants which optimize MSE, , and , in addition to an adversarial loss , ENet (“PAT” variant), Deng (), and Mechrez et al. .
Interestingly, the same pattern is observed in all plots: (i) The lower left corner is blank, revealing an unattainable region in the perception-distortion plane. (ii) In proximity of this blank region, NR and FR metrics are anti-correlated, indicating a tradeoff between perception and distortion. Notice that the tradeoff exists even for the IFC, VIF and VGG2,2 measures, which are considered to capture visual quality better than MSE and SSIM.
Figure 10 depicts the outputs of several algorithms lying closest to the perception-distortion bound in the IFC graph in Fig. 9. While the images are ordered from low to high distortion (according to IFC), their perceptual quality clearly improves from left to right.
Both FR and NR measures are commonly validated by calculating their correlation with human opinion scores, based on the assumption that both should be correlated with perceptual quality. However, as Fig. 11 shows, while FR measures can be well-correlated with perceptual quality when distant from the unattainable region, this is clearly not the case when approaching the perception-distortion bound. In particular, all tested FR methods are inconsistent with human opinion scores which found the SRGAN to be superb in terms of perceptual quality , while NR methods successfully determine this. We conclude that image restoration algorithms should always be evaluated by a pair of NR and FR metrics, constituting a reliable, reproducible and simple method for comparison, which accounts for both perceptual quality and distortion. This evaluation method was demonstrated and validated by a human opinion study in the 2018 PIRM super-resolution challenge .
Up until 2016, SR algorithms occupied only the upper-left section of the perception-distortion plane. Nowadays, emerging techniques are exploring new regions in this plane. The SRGAN, ENet, Deng, Johnson et al. and Mechrez et al. methods are the first (to our knowledge) to populate the high perceptual quality region. In the near future we will most likely witness continued efforts to approach the perception-distortion bound, not only in the low-distortion region, but throughout the entire plane.
Conclusion
We proved and demonstrated the counter-intuitive phenomenon that distortion and perceptual quality are at odds with each other. Namely, the lower the distortion of an algorithm, the more its distribution must deviate from the statistics of natural scenes. We showed empirically that this tradeoff exists for many popular distortion measures, including those considered to be well-correlated with human perception. Therefore, any distortion measure alone, is unsuitable for assessing image restoration methods. Our novel methodology utilizes a pair of NR and FR metrics to place each algorithm on the perception-distortion plane, facilitating a more informative comparison of image restoration methods.
Acknowledgements This research was supported by the Israel Science Foundation (grant no. 852/17), and by the Technion Ollendorff Minerva Center.
References
Appendix A Real-vs.-fake user studies and hypothesis testing
We assume the setting where an observer is shown a real image (a draw from ) or an algorithm output (a draw from ), with a prior probability of each. The task is to identify which distribution the image was drawn from ( or ) with maximal probability of success. This is the setting of the Bayesian hypothesis testing problem, for which the maximum a-posteriori (MAP) decision rule minimizes the probability of error (see Section 1 in ). When there are two possible hypotheses with equal probabilities (as in our setting), the relation between the probability of error and the total-variation distance between and in (1) can be easily derived (see Section 2 in ).
Appendix B The MMSE and MAP examples of Sec. 3
Sections 3.1 and 3.2 exemplify that the MSE and the loss are not distribution preserving in the setting of estimating a discrete random variable (vector) from its noisy version , where is independent of . Since the conditional distribution of given is , the MMSE estimator is given by
In the example of Fig. 4, is a binary image comprising blocks chosen uniformly at random from a finite database. Since the noise is i.i.d., each block of can be denoised separately, both in the case of the MSE criterion and in the case of MAP. For each block, we have for the non-blank images and for the blank image.
In the trinary example (5), we calculate the distribution of the MMSE estimate (Fig. 3) by
where the inverse of (see (6)) and its derivative are calculated numerically, and with and of (5).
Appendix C Proof of Theorem 1
We will show that a stably distribution preserving optimal estimator is necessarily unique. At the same time, we will show that a non-invertible degradation implies that this optimal estimator is non-unique. Specifically, we use the following definitions.
We say that the optimal estimator is not unique if there exist two estimators, and that minimize the mean distortion (3) and differ from one another in the sense that
The outline of the proof of Theorem 1 will be as follows:
In Lemma 1 we will show that if the distortion measure is stably distribution preserving, then the optimal estimator is uniquely defined by .
In Lemma 2 we will show that if the estimator defined by is an optimal estimator and the degradation is non-invertible, then the optimal estimator is non-unique.
This leads to a contradiction, proving that there does not exist a stably distribution preserving distortion metric if the degradation in non-invertible.
If the distortion measure is stably distribution preserving at , then the optimal estimator that minimizes the mean distortion (3) is uniquely defined by .
We start by noting that the optimal estimator depends only on and not on . Indeed, since and are independent given , the mean distortion can be written as
Therefore, the optimal is that which minimizes for each . Since depends only on , the optimal estimator depends only on .
where we used the assumption that . Similarly, the distribution of has changed to
This equality can hold for every perturbation only if , completing the proof.
Notice that this also proves that the optimal estimator is unique (under the stably distribution preserving assumption), as we demonstrated that only minimizes the mean distortion. ∎
If the degradation is non-invertible, and the estimator defined by is an optimal estimator, then the optimal estimator is non-unique.
Now, since is an optimal estimator, must minimize for each (see proof of Lemma 1). This means that for any , the conditional must assign positive probability only to in the set of minima . We conclude that for every . This implies that any other estimator that assigns zero probability to for every , is also optimal.
Let be non-empty disjoint sets such that . Now, define two estimators, such that only for , and only for , for every . Both are optimal estimators (as they only assign positive probability to ). Yet, these two estimators have conditional distributions with disjoint supports for every , and thus . Therefore, by Definition 7, the optimal estimator is non-unique. ∎
Now, let us assume to the contrary that defines a non-invertible degradation, and that the distortion function is stably distribution preserving at . By Lemma 1, the optimal estimator is uniquely defined by . But now according to Lemma 2, since the degradation is non-invertible, the optimal estimator is non-unique, leading to a contradiction.
Appendix D Derivation of Example 1
Since , it is a zero-mean Gaussian random variable. Now, the Kullback-Leibler distance between two zero-mean normal distributions is given by
and the MSE between and is given by
Substituting and , we obtain that and , so that
Notice that is symmetric, and (see Fig. 12). Thus, for any negative , there always exists a positive with which is the same and the MSE is not larger. Therefore, without loss of generality, we focus on the range .
For the constraint set of is empty, and there is no solution to (33). For , the constraint is satisfied for , where
For , the optimal (and only possible) is
For , monotonically increases with , broadening the constraint set. The objective monotonically decreases with in the range (see Fig. 12 and the mathematical justification below). Thus, for , the optimal is always the largest possible , which is , where is defined by (see Fig. 12). For , the optimal is , which achieves the global minimum . The closed form solution is therefore given by
To justify the monotonicity of in the range , notice that for ,
which is negative for .
Appendix E Proof of Theorem 2
The proof of Theorem 2 follows closely that of the rate-distortion theorem from information theory . The value is the minimal distance over a constraint set whose size does not decrease with . This implies that the function is non-increasing in . Now, to prove the convexity of , we will show that
for all (see Fig. 13). First, by definition, the left hand side of (38) can be written as
where and are the estimators defined by
Since is convex in its second argument,
because is in the constraint set. Below, we show that
Therefore, since is non-increasing in , we have that
Combining (39), (42), (E) and (46) proves (38), thus demonstrating that is convex.
where the second and fourth transitions are according to the law of total expectation and the third transition is justified by
Here we used (43) and the fact that and are independent given , and similarly for the pairs and .
Appendix F Proof of Theorem 3
The estimator of (14) attains perfect perceptual quality since
where we used the law of total expectation and the fact that given , and are independent and identically distributed. The MSE of is therefore
where the second equality is due to (F) and (51), and the third equality is due to the orthogonality principle. We thus established that is a distribution preserving estimator whose MSE is precisely twice the MSE of the MMSE estimator. This implies that
Appendix G WGAN architecture and training details (Sec. 5)
Appendix H Super-resolution evaluation details (Sec. 6) and additional comparisons
The no-reference (NR) and full-reference (FR) methods BRISQUE, BLIINDS-II, NIQE, SSIM, MS-SSIM, IFC and VIF were obtained from the LIVE laboratory websitehttp://live.ece.utexas.edu/research/Quality/index.htm, the NR method of Ma et al. was obtained from the project webpagehttps://github.com/chaoma99/sr-metric, and the pretrained VGG-19 network was obtained through the PyTorch torchvision packagehttp://pytorch.org/docs/master/torchvision/index.html. The low-resolution images were obtained by factor 4 downsampling with a bicubic kernel. The super-resolution results on the BSD100 dataset of the SRGAN and SRResNet variants were obtained onlinehttps://twitter.box.com/s/lcue6vlrd01ljkdtdkhmfvk7vtjhetog, and the results of EDSR, Deng, Johnson et al. and Mechrez et al. were kindly provided by the authors. The algorithms for testing the other SR methods were obtained online: A+http://www.vision.ee.ethz.ch/~timofter/ACCV2014_ID820_SUPPLEMENTARY/, SRCNNhttp://mmlab.ie.cuhk.edu.hk/projects/SRCNN.html, SelfExhttps://github.com/jbhuang0604/SelfExSR, VDSRhttp://cv.snu.ac.kr/research/VDSR/, LapSRNhttps://github.com/phoenix104104/LapSRN, Bae et al. https://github.com/iorism/CNN and ENethttps://webdav.tue.mpg.de/pixel/enhancenet/. All NR and FR metrics and all SR algorithms were used with the default parameters and models. In the paper, we reported comparisons on the y-channel (except for the measure). In the supplementary material, we report results with additional NR metrics on the y-channel, as well as results on color images. When comparing color images, for SR algorithms which treat the y-channel alone, the Cb and Cr channels are upsampled by bicubic interpolation.
The general pattern appearing in Fig. 9 will appear for any NR method which accurately predicts the perceptual quality of images. We show here three additional popular NR methods: BRISQUE , BLIINDS-II and the recent measure by Ma et al. in Figs. 14, 15, 16, where the same conclusions as for NIQE (see Sec. 6) are apparent. The same pattern appears for RGB images as well, as shown in Figs. 17, 18. Note that the perceptual quality of Johnson et al. and SRResNet-VGG2,2 is inconsistent between NR metrics, likely due to varying sensitivity to the cross-hatch pattern artifacts which are present in these method’s outputs. For this reason, Johnson et al. does not appear in the NIQE plots, as its NIQE score is (far off the plots).