Rethinking Lossy Compression: The Rate-Distortion-Perception Tradeoff
Yochai Blau, Tomer Michaeli
Introduction
Lossy compression techniques are ubiquitous in the modern-day digital world, and are regularly used for communicating and storing images, video and audio. In recent years, lossy compression is seeing a surge of research, due in part to the advancements in deep learning and their application in this domain (Toderici et al., 2016, 2017; Ballé et al., 2016, 2017, 2018; Agustsson et al., 2017, 2018; Rippel & Bourdev, 2017; Minnen et al., 2018; Li et al., 2018; Mentzer et al., 2018; Johnston et al., 2018; Galteri et al., 2017; Tschannen et al., 2018; Santurkar et al., 2018; Rott Shaham & Michaeli, 2018). The theoretical foundations of lossy compression are rooted in Shannon’s seminal work on rate-distortion theory (Shannon, 1959), which analyzes the fundamental tradeoff between the bit rate used for representing data, and the distortion incurred when reconstructing the data from its compressed representation (Cover & Thomas, 2012).
The premise in rate-distortion theory is that reduced distortion is a desired property. However, recent works demonstrate that minimizing distortion alone does not necessarily drive the decoded signals to have good perceptual quality. For example, incorporating generative adversarial type losses has been shown to lead to significantly better perceptual quality, but at the cost of increased distortion (Tschannen et al., 2018; Agustsson et al., 2018; Santurkar et al., 2018). This behavior has also been studied in the context of signal restoration (Blau & Michaeli, 2018), where it was shown that minimizing distortion causes the distribution of restored signals to deviate from that of the ground-truth signals (indicating worse perceptual quality). In light of this understanding, it is natural to seek for a generalized rate-distortion theory, which also accounts for perception. In particular, it is of key importance to understand how the best achievable rate depends not only on the distortion, but also on the perceptual quality of the algorithm. A preliminary attempt to incorporate perceptual quality into rate-distortion theory was briefly reported in (Matsumoto, 2018a, b). Yet, no theoretical characterization nor practical demonstration of its effect on the rate-distortion tradeoff was presented.
In this paper, we adopt the mathematical definition of perceptual quality used in (Blau & Michaeli, 2018), and prove that there is a triple tradeoff between rate, distortion and perception. Our key observation is that the rate-distortion function elevates as the perceptual quality is enforced to be higher (see Fig. 1). In other words, to obtain good perceptual quality, it is necessary to make a sacrifice in either the distortion or the rate of the algorithm.
Our analysis is based on the definition of a rate-distortion-perception function , which characterizes the minimal achievable rate for any given distortion and perception index . We begin by deriving a closed form for this function in the classical case study of a Bernoulli source, a simple example which nonetheless nicely illustrates the typical behavior of the tradeoff. We then prove several general properties of , showing that it is monotone and convex for any full-reference distortion measure (under minor assumptions), and that there is a range of values for which it necessarily does not coincide with the traditional rate-distortion function. For the specific case of the squared-error distortion, we also provide an upper bound on the increase in distortion that has to be incurred in order to achieve perfect perceptual quality, at any given rate.
Our observations have important implications for the design and evaluation of practical compression methods. In particular, they suggest that comparing between algorithms only in terms of their rate-distortion curves can be misleading. We demonstrate this in the context of image compression using a toy MNIST example, by systematically exploring the visual effect of improvement in each of the three properties (rate, distortion, perception) on the expense of the others. We do this by training an encoder-decoder net utilizing a generative model, similarly to (Tschannen et al., 2018; Agustsson et al., 2018). As we show, the phenomena we discuss are dominant at low bit rates, where the classical approach of optimizing distortion alone leads to unacceptable perceptual quality. This is perhaps not surprising when using the MSE distortion, which is known to be inconsistent with human perception. But our theory shows that every distortion measure (excluding pathological cases) must have a tradeoff with perceptual quality. This includes e.g., the popular SSIM/MS-SSIM (Wang et al., 2003, 2004), the distance between deep features (Johnson et al., 2016), and any other full reference criterion. To illustrate this, we repeat our toy experiment with the distortion measure of (Johnson et al., 2016), which has been used as a means for enhancing perceptual quality in low-level vision tasks (Ledig et al., 2017). As we show, minimizing this distortion does not lead to good perceptual quality at low bit rates, just like our theory predicts. Moreover, when enforcing high perceptual quality, this distortion rather increases.
Background
Rate-distortion theory analyzes the fundamental tradeoff between the rate (bits per sample) used for representing samples from a data source , and the expected distortion incurred in decoding those samples from their compressed representations. Formally, the relation between the input and output of an encoder-decoder pair, is a (possibly stochastic) mapping defined by some conditional distribution , as visualized in Fig. 2. The expected distortion of the decoded signals is thus defined as
A key result in rate-distortion theory states that for an iid source , if the expected distortion is bounded by , then the lowest achievable rate is characterized by the (information) rate-distortion function
where denotes mutual information (Cover & Thomas, 2012). Closed form expressions for the rate-distortion function are known for only a few source distributions and under quite simple distortion measures (e.g., squared error or Hamming distance). However several general properties of this function are known, including that it is always monotonically non-increasing and convex.
2 Perceptual Quality
The perceptual quality of an output sample refers to the extent to which it is perceived by humans as a valid (natural) sample, regardless of its similarity to the input . In various domains, perceptual quality has been associated with the deviation of the distribution of output signals from the distribution of natural signals, which, as discussed in (Blau & Michaeli, 2018), is linked to the common practice of quantifying perceptual quality via real-vs.-fake user studies (Isola et al., 2017; Salimans et al., 2016; Zhang et al., 2016; Denton et al., 2015). In particular, deviation from natural scene statistics is the basis for many no-reference image quality measures (Mittal et al., 2013, 2012; Wang & Simoncelli, 2005), which have been shown to correlate well with human opinion scores. It is also the principle underlying GAN-based image restoration schemes, which achieve enhanced perceptual quality by directly minimizing some divergence (Ledig et al., 2017; Pathak et al., 2016; Isola et al., 2017; Wang et al., 2018). Based on these works, and following (Blau & Michaeli, 2018), we define the perceptual quality index (lower is better) of an algorithm as
where is some divergence between distributionsWe assume that . (e.g., Kulback-Leibler, Wasserstein, etc.). Note that the divergence function which best relates to human perception is a subject of ongoing research. Yet, our results below hold for (nearly) any divergence.
Obviously, perceptual quality, as defined above, is very different from distortion. In particular, minimizing the perceptual quality index does not necessarily lead to low distortion. For example, if the decoder disregards the input, and outputs random samples from the source distribution , it will achieve perfect perceptual quality but very poor distortion. It turns out that this is true also in the other direction. That is, minimizing distortion does not necessarily lead to good perceptual quality. This observation has been studied in (Blau & Michaeli, 2018) in the specific context of signal restoration (e.g. denoising, super-resolution). In particular, perception and distortion are fundamentally at odds with each other (for non-invertible degradations), in the sense that optimizing one always comes on the expense of the other. This behavior, coined the perception-distortion tradeoff, was shown to hold true for any distortion measure.
The Rate-Distortion-Perception Tradeoff
Since both perceptual quality and distortion are typically important, here we extend the rate-distortion function (2) to take into account the perception indexSimilarly to (2), in (1) lower bounds the best achievable rate for an iid source (see Supplementary Material). We do not prove achievability of in general. However for the MSE distortion, we show an achievable upper bound (see Theorem 2). (3).
The (information) rate-distortion-perception function is defined as
Unfortunately, closed form solutions for (1) are even harder to obtain than for (2). Yet, one notable exception is the classical case study of a binary source, as we show next. While of limited applicability, this example illustrates the typical behavior of (1), which we analyze in Sec. 3.2.
Consider the problem of encoding a binary source , where the decoder’s output is also constrained to be binary. Let us take the distortion measure to be the Hamming distance, and the perception indexThe term “perception” is somewhat inappropriate for a Bernoulli source, as it is not perceived by humans (contrary to images, audio). Yet, we keep this terminology here for consistency. to be the total-variation (TV) distance . Without loss of generality, we assume that . When perception is not constrained (i.e., ), the solution to (1) reduces to the rate-distortion function (2) of a binary source, which is known to be given by
where is the entropy of a Bernoulli random variable with probability (Cover & Thomas, 2012).
In the Supplementary Material, we derive the solution for arbitrary . It turns out that as long as the perceptual quality constraint is sufficiently loose, the solution remains the same. However, when , the perception constraint in (1) becomes active whenever the distortion constraint is loose enough, from which point the function departs from . Specifically, for , we have
where and denotes the entropy of a ternary random variable with probabilities , , . Here, , , and , where and .
Figure 3 plots as a function of for several values of . As can be seen, at , all the curves merge. This is because at this point (lossless compression), so that , and thus the perceptual quality is perfect. Yet, as the allowed distortion grows larger, the curves depart. This illustrates that achieving the classical rate-distortion curve (black dashed line) does not generally lead to good perceptual quality. The more stringent our prescribed perceptual quality constraint (lower ), the more the rate-distortion curve elevates (colored curves). In particular, the tradeoff becomes severe at the low bit rate regime, where good perceptual quality comes at the cost of a significantly higher distortion and/or bit rate. Notice that it is possible to achieve perfect perceptual quality at every rate (blue curve) by compromising the distortion to some extent. In Sec. 3.2 we provide an upper-bound on the increase in distortion required for obtaining perfect perceptual quality.
While Fig. 3 displays cross-sections of along rate-distortion planes, in Fig. 4 we plot as a surface in 3 dimensions, as well as its cross-sections along the other planes. The equi-rate level sets shown on the surface in Fig. 4, provide another visualization for the phenomenon described above. That is, at high bit-rates, it is possible to achieve good perceptual quality (low ) without a significant sacrifice in the distortion . However, as the bit-rate becomes lower, the equi-rate level sets substantially curve towards the low values, illuminating the exacerbation in the tradeoff between distortion and perception in this regime. Figure 4 provides an additional viewpoint, by showing perception-distortion curves for different bit rates. Notice again that the tradeoff between distortion and perceptual quality becomes stronger at low bit-rates. Finally, Fig. 4 shows the somewhat counter-intuitive tradeoff between rate and perceptual quality as a function of distortion. Specifically, we see that at every constant distortion level, the perceptual quality can be improved by increasing the rate.
2 Theoretical Properties
For general source distributions, it is usually impossible to solve (1) analytically. However, it turns out that the behavior we saw for a Bernoulli source is quite typical. We next prove several general properties of the function (1), which hold under rather mild assumptions. Specifically, we assume:
A1 The divergence in (1) is convex in its second argument. That is, for any and for any three distributions ,
Assumption A1 is not very limiting. For instance, any -divergence (e.g. KL, TV, Hellinger, ) as well as the Renyi divergence, satisfies this assumption (Csiszár et al., 2004; Van Erven & Harremos, 2014). Assumption A2 holds in any setting where the mean distance between a “valid” signal and all other “valid” signals is not constantA valid signal is any . Also, we use “distance” here for clarity, although is not necessarily a metric.. In particular, it holds for any distortion function with a unique minimizer, such as the squared-error distortion and the SSIM index (under some assumptions) (Brunet, 2012). Using these assumptions, we are able to qualitatively characterize the general shape of the function .
The rate-distortion-perception function (1):
is monotonically non-increasing in and ;
satisfies if A2 holds.
The proof of Theorem 1 can be found in the Supplementary Material. Note that when assumption A2 holds, properties 1 and 3 indicate that there exists some for which , showing that the rate-distortion curve necessarily elevates when constraining for perfect perceptual quality. In any case, assumption A2 is a sufficient condition for property 3, so that even if it does not hold, this does not necessarily imply that .
How much does the rate-distortion curve elevate when constraining for perfect perceptual quality? The next theorem upper-bounds this elevation for the MSE distortion (see proof in the Supplementary Material).
When using the squared-error distortion, the function (rate-distortion at perfect perceptual quality) is bounded by
Theorem 2 shows that it is possible to attain perfect perceptual quality without increasing the rate, by sacrificing no more than a -fold increase in the mean squared-error (MSE). More specifically, attaining perfect perceptual quality at distortion does not require a higher bit rate than that necessary for compression at distortion with no perceptual quality constraint. This is illustrated in Fig. 5, where the perfect-quality curve shown in blue is bounded by the scaled version of Shannon’s unconstrained quality curve shown as a black dashed line. In image restoration scenarios, such a -fold increase in the MSE (dB decrease in PSNR) has been shown to enable a substantial improvement in perceptual quality by practical algorithms (Blau et al., 2018; Ledig et al., 2017). Note that this bound is generally not tight. Thus, in some settings, perfect perceptual quality can be obtained with an even smaller increase in distortion.
Experimental Illustration
We now turn to demonstrate the visual implications of the rate-distortion-perception tradeoff in lossy image compression on a toy MNIST example. We make no attempt to propose a new state-of-the-art compression method. Our sole goal is to systematically explore the effect of the balance between rate, distortion, and perception. To this end, we utilize a net-based encoder-decoder pair trained in an end-to-end fashion, similarly to recent works. By tuning the influence of each of the different terms of the loss, we can easily control the balance between these three quantities.
More concretely, we use an encoder and a decoder , both parametrized by deep neural nets (DNNs). The encoder maps the input into a latent feature vector , whose entries are then uniformly quantized to levels to obtain the representation . The decoder outputs a reconstruction . To enable back-propagation through the quantizer, we use the differentiable relaxation of (Mentzer et al., 2018). Note that this relaxation affects only the gradient computation through the quantizer during back-propagation, but not the forward-pass “hard” quantization.
As in recent perceptual-quality driven lossy compression schemes (Tschannen et al., 2018; Agustsson et al., 2018), the rate is controlled by the dimension of the encoder’s output , and the number of levels used for quantizing each of its entries, such that . Note that this only upper-bounds the best achievable rate, as lossless compression of would potentially further reduce the representation’s size. However, it significantly simplifies the scheme, and was found to be only slightly sub-optimal (Agustsson et al., 2018).
For any fixed rate, we train the encoder-decoder to minimize a loss comprising a weighted combination of the expected distortion and the perception index,
where in our specific case of a Wasserstein GAN (Arjovsky et al., 2017), denotes the class of bounded -Lipschitz functions. As usual, all expectations are replaced by sample means, the constraint is replaced by a gradient penalty (Gulrajani et al., 2017), and the loss is minimized by alternating between minimization w.r.t. while holding fixed and maximization w.r.t. while holding fixed.
To achieve good perceptual quality, especially at low rates, it is essential that the decoder be stochastic (Tschannen et al., 2018). This is commonly carried out by an additional random noise input. Yet, deep generative models in the conditional setting tend to ignore this type of stochasticity (Zhu et al., 2017a, b; Mathieu et al., 2016). Tschannen et al. (2018) remedy this by applying a two-stage training scheme, which indeed promotes the use of stochasticity within the decoder, but can lead to sub-optimal results. Here, instead of concatenating a noise vector to the encoder’s output , we add it, so that the decoder in fact operators on the noisy representation . This does not lead to loss of information, as the noise is drawn from a uniform distribution , with smaller than the quantization bin size. Thus, different coded representations do not “mix-up”, and can always be distinguished from one another. This scheme urges the decoder to utilize the stochastic input, while allowing end-to-end training in a one-step manner.
We begin by experimenting with the squared-error distortion . We train 98 encoder-decoder pairs on the MNIST handwritten digit dataset (LeCun et al., 1998), while varying the encoder’s output dimension and number of quantization levels to control the rate , and the tuning coefficient to achieve different balances between distortion and perceptual quality. A list of all combinations of used, along with all other training details can be found in the Supplementary Material.
The left side of Fig. 6 plots the 98 trained encoder-decoder pairs on the rate-distortion plane, with the perceptual quality indicated by color coding and rate measured in bits per digit. The perceptual quality is quantified by the final discriminator loss, where are the test samples., which approximates the Wasserstein distance . We plot an approximation of Shannon’s rate-distortion function (obtained with ), and two additional rate-distortion curves with (approximately) constant perceptual qualityWe plot a smoothing spline calculated over the set of points which satisfy the constraint and have the minimal distortion among all points with the same rate.. As can be seen, the rate-distortion curve elevates when constraining the perceptual quality to be good. This demonstrates once again that we can improve the perceptual quality w.r.t. that obtained on Shannon’s rate-distortion curve, yet this must come at the cost of a higher rate and/or distortion. Notice that the perception index is not constant along Shannon’s function; it increases (worse quality) towards lower bit-rates.
On the right side of Fig. 6, we depict the outputs of encoder-decoder pairs along Shannon’s rate-distortion function, and along the two equi-perception curves shown on the left. It can be seen that as the rate decreases, the perceptual quality of the reconstructions along Shannon’s function degrades. However, this is avoided when constraining the perceptual quality, which results in visually pleasing reconstructions even at extremely low bit-rates. Notice that this increased perceptual quality does not imply increased accuracy, as at low bit rates (e.g., bits), most reconstructions fail to preserve even the identity of the digit. Yet, while the encoder-decoder pairs on Shannon’s rate-distortion curve are more accurate on average, no doubt that the perceptually-constrained encoder-decoder pairs are favorable in terms of perceptual quality. Also, notice that at a rate of bits, the outputs of the perceptually-constrained encoder-decoder pairs are all distinct, even though there are only code words (as can be seen for Shannon’s encoder-decoder). This shows that the decoder effectively utilizes the noise.
Figure 7 depicts the function in 3-dimensions, as well as its cross sections along the other axis aligned planes. In Fig. 7, the curved equi-rate lines show the tradeoff between distortion and perceptual quality. This is also apparent in Fig. 7, which shows cross sections along perception-distortion planes at different rates. As can be seen, the tradeoff becomes stronger at low bit-rates. Figure 7 shows the counter-intuitive tradeoff between rate and perception. That is, at constant distortion, the perceptual quality can be improved by increasing the rate.
2 Advanced Distortion Measures
The peak-signal-to-noise ratio (PSNR), which is a rescaling of the MSE, is still the most common quality measure in image compression. Yet, it is well-known to be inadequate for quantifying distortion as perceived by humans (Wang & Bovik, 2009). Over the past decades, there has been a constant search for better distortion criteria, ranging from the simple SSIM/MS-SSIM (Wang et al., 2003, 2004) to the recently popular deep-feature based distortion (Johnson et al., 2016; Zhang et al., 2018). Interestingly, the perceptual quality along Shannon’s classical rate-distortion function is not perfect for nearly any distortion measure (see property 3 in Theorem 1). This implies that perfect perceptual quality cannot be achieved by merely switching to more advanced distortion criteria, but rather requires directly optimizing the perception index (e.g. using GAN-based schemes). This is not to say that the function is the same for all distortion measures. The strength of the tradeoff can certainly decrease for distortion criteria which capture more semantic similarities (Blau & Michaeli, 2018).
We demonstrate this by repeating the experiment of Sec. 4.1, while replacing the squared-error distortion by the deep-feature based distortion of (Ledig et al., 2017), i.e.,
where is the output of an intermediate DNN layer for input . Here we take the second conv-layer output of a -layer DNN, which we pre-trained to achieve over classification accuracy on the MNIST test set. All training details appear in the Supplementary Material.
Figure 8 plots 98 encoder-decoder pairs on the rate-distortion plane, trained exactly as in Fig. 6, but this time with the loss (11) instead of MSE. As can be seen, here too the rate-distortion curves elevate when constraining the perceptual quality, demonstrating that the use of advanced distortion measures does not eliminate the tradeoff. From the decoded outputs, however, it is evident that the tradeoff here is somewhat weaker, as minimizing distortion alone (Shannon’s) appears a bit more visually pleasing compared to Fig. 6 (though still with reduced variability and more blur than the perception constrained reconstruction).
3 Related Work
Our theoretical analysis and experimental validation help explain some of the observations reported in the recent literature. Specifically, a lot of research efforts have been devoted to optimizing the rate-distortion function (2) using deep nets (Toderici et al., 2016, 2017; Agustsson et al., 2017; Ballé et al., 2017; Minnen et al., 2018; Li et al., 2018). Some papers explicitly targeted high perceptual quality. One line of works did so by choosing the distortion criterion to be some advanced full-reference measure, like SSIM/MS-SSIM (Ballé et al., 2018; Mentzer et al., 2018; Johnston et al., 2018), normalized Laplacian pyramid (Ballé et al., 2016) and deformation-aware sum of squared differences (DASSD) (Rott Shaham & Michaeli, 2018). While beneficial, these methods could not demonstrate high perceptual quality at very low bit rates, which aligns with our theory. Another line of works incorporated generative models, which explicitly encourage the distribution of outputs to be similar to that of natural images (decreasing the divergence in (3)). This was done on an image patch level (Rippel & Bourdev, 2017), on reduced-size (thumbnail) images (Tschannen et al., 2018; Santurkar et al., 2018), on a full-image scale (Agustsson et al., 2018), and as a post-processing step (Galteri et al., 2017). In particular, Tschannen et al. (2018) propose a practical method for distribution-preserving compression ( in our terminology). These methods managed to obtain impressive perceptual quality at very low bit rates, but not without a substantial sacrifice in distortion, as predicted by our theory. Finally, we note that rate-distortion analysis (with a specific distortion) has also been used in the context of generative models (Alemi et al., 2018), which target (i.e., ). Our results hold for arbitrary distortions and arbitrary .
Conclusion
We proved that in lossy compression, perceptual quality is at odds with rate and distortion. Specifically, any attempt to keep the statistics of decoded signals similar to that of source signals, will result in a higher distortion or rate. We characterized the triple tradeoff between rate, distortion and perception, and empirically illustrated its manifestation in image compression. Our observations suggest that comparing methods based on their rate-distortion curves alone may be misleading. A more informative evaluation must also include some (no-reference) perceptual quality measure.
Acknowledgements
This research was supported in part by the Israel Science Foundation (grant no. 852/17) and by the Ollendorf Foundation.
References
Appendix A Proof of Theorem 1
The proof of this theorem follows closely that of its rate-distortion analogue (Cover & Thomas (2012), 2nd ed., p. 316).
The value is the minimal mutual information over a constraint set whose size increases with and . This implies that the function is non-increasing in and .
Convexity
Here, we assume that A1 holds. That is, the divergence in (1) is convex in its second argument, so that for any ,
To prove the convexity of , we will show that
for all . First, by definition, the left hand side of (13) can be written as
where and are defined by
Since is convex in for a fixed (Cover & Thomas (2012), 2nd ed., p. 33),
because is in the constraint set. The divergence is assumed to be convex in the second argument, thus
where (a) and (c) are according to the law of total expectation, and (b) is by (18). Therefore, since is non-increasing in and , we have from (A) and (A) that
Combining (14), (17), (19) and (22) proves (13), thus proving that is convex.
Dependence on the perceptual quality
Notice that since in this case, and are independent, so that . Therefore
Clearly, the which minimizes (A) cannot assign positive probability outside the set where attains its minimal value. Namely, the support of must be contained in the set defined by
But since our encoder-decoder pair achieves perfect perceptual quality, i.e. , this implies that , contradicting Assumption A2.
Appendix B Proof of Theorem 2
Appendix C Perception aware lossy compression of a memoryless stationary source
We now prove that when compressing a memoryless stationary source with average distortion and average perception index , the rate is lower bounded by . This proof follows closely that of its rate-distortion analogue (Cover & Thomas (2012), 2nd ed., p. 316).
Assume a memoryless stationary source. Given a source sequence comprising i.i.d. variables with distribution , the encoder constructs an encoded representation with rate as . The decoder outputs an estimate of as . We are interested in the the average distortion of the reconstructions, , and in their average perceptual quality, . Assume that
where (a) is since the size of the range of is , (b) is since , (c) is from the data-processing inequality, (d) is since are independent, (e) is from the chain rule of entropy, (f) is since conditioning reduces entropy, (g) is from the definition of in (1), (h) is from the convexity of (see Theorem 1) and Jensen’s inequality, and (i) is from (31) and the fact that is non-increasing in (see Theorem 1). This proves that the rate of any encoder-decoder pair having average distortion and average perceptual quality , is lower-bounded by , the rate-distortion-perception function evaluated at .
To prove that the rate-distortion-perception function describes the optimal rate at distortion level and perceptual quality , we would also have to prove that is achievable, which we leave for future work. Yet, the proof that lower-bounds the rate is sufficient for concluding that a tradeoff between rate, distortion and perception necessarily exists. Specifically, in Theorem 1 we prove that (subject to assumptions) the rate-distortion curve elevates when constraining for perceptual quality, i.e. . Now, Shannon’s rate-distortion curve is known to be achievable (Cover & Thomas (2012), 2nd ed., p. 318) and thus describes the optimal rate when not constraining the perceptual quality. As shown above, lower-bounds the rate when constraining for perfect perceptual quality. Combining these, we get that , indicating that constraining for perceptual quality necessarily leads to an increase in rate (for constant distortion level), thus illustrating the rate-distortion-perception tradeoff.
Appendix D Derivation of the rate-distortion-perception function R(D,P)𝑅𝐷𝑃R(D,P) of a Bernoulli source
Assume that with . We seek a conditional distribution , which we parameterize by as
that solves the rate-distortion-perception problem
Here we concentrate on the case where is the Hamming distance, and is the total-variation (TV) divergence. The mutual information term is given by
The function for Shannon’s classic rate-distortion problem is given by (see (Cover & Thomas, 2012), 2nd ed., p. 308)
where denotes the binary entropy . This optimal solution is obtained by setting the parameters to
Solution for finite P𝑃P and I(X,X^)>0𝐼𝑋^𝑋0I(X,\hat{X})>0
Therefore, the constraint is satisfied when
Below, we show that the lower constraint of (43) is never active (see J1). The upper constraint is obviously active only when of (40) does not satisfy the upper bound in (43), which happens when
Therefore, when the solution is independent of and is given by (39). When , the constraint is active, the upper constraint of (43) is active, and thus
and by substituting into (41) we also get
Note that in (44) we assumed , below we will justify that this is always the case in this region (see J2).
Now, substituting from (45), (46) back into (36) we get
where and . This can be further simplified to obtain
where is the entropy of a ternary random variable (taking values in a three element alphabet) with probabilities .
Solution for finite P𝑃P and I(X,X^)=0𝐼𝑋^𝑋0I(X,\hat{X})=0
The function is non-increasing in (see Theorem 1), and will reach for since in this case and are independent. From (45) and (46), this happens when
From this point onward, the solution is fixed, as mutual information is non-negative and we cannot further decrease the objective of (35).
Overall solution
Putting all the pieces together, the overall solution for is
where and are defined in (44) and (49), respectively. For , the solution is independent of and is given by the solution to Shannon’s classic rate-distortion curve for a Bernoulli source in (39) (see justification in J3 below).
Additional justifications
The solution in (40) does not satisfy this lower constraint of (43) when
When this happens for , which never occurs as . When this happens for . However, since for all (which is our assumption), the upper constraint of (43) will always become active before the lower constraint.
J2
Taking the derivative of with respect to we obtain
which is non-negative since . Thus, is increasing in for all , and its largest value in the range , which is , is obtained at . Thus, in the region where , it is ensured that .
J3
Taking the derivative of with respect to we obtain
which is non-negative for , thus is non-decreasing in . It is easy to see from (49) that in non-increasing in (for ). Thus, for a single , which is . For any , there are no satisfying .
Appendix E Architecture and training parameters for the experiments in Sec. 4
The architecture of the encoder, decoder and discriminator nets used for compressing (and decompressing) the MNIST images in Sec. 4 is detailed in Table 1. The optimization objective is given in (10), where is the squared-error distortion in Sec. 4.1, and a combination of the squared-error and the “perceptual loss” of Johnson et al. (2016) in Sec. 4.2. The encoder output dimension , the number of quantization levels , and values of the tradeoff coefficient in (10) used for training the encoder-decoder pairs appear in Table 2. The distortion term in (10) was also multiplied by a constant factor of for the MSE term (in Sec. 4.1 and Sec. 4.2) and factor of for the perceptual loss (in Sec. 4.2). For each and , an encodoer-decoder with (only distortion, no adversarial loss) was trained for 25 epochs. The other encoder-decoder pairs with continued training from this point for another 25 epochs. The ADAM optimizer was used with . Batch size was . Initial learning rates were for the encoder-decoder/discriminator updates in Sec. 4.1, and for the encoder-decoder/discriminator updates in Sec. 4.2. These learning rates decreased by after 20 epochs. The convolutional/transposed-convolutional layers filter size (in the decoder and discriminator) was always , except for the last convolutional layer in the decoder where the filter size was . No padding was used in the decoder, and a padding of was used in each convolutional layer of the discriminator.
The quantization layer (last encoder layer) follows Mentzer et al. (2018). Here, the bin centers are fixed and evenly spaced in the interval $z_{i}i\hat{z}_{i}\hat{z}_{i}=\arg\min_{c_{j}}\|z_{i}-c_{j}\|$. To compute the gradients in the backward pass, we use a differential “soft” assignment
where we use . Uniformly distributed noise is added to the encoder output before it is passed on to the decoder, with .