PULSE: Self-Supervised Photo Upsampling via Latent Space Exploration of Generative Models

Sachit Menon, Alexandru Damian, Shijia Hu, Nikhil Ravi, Cynthia Rudin

Introduction

In this work, we aim to transform blurry, low-resolution images into sharp, realistic, high-resolution images. Here, we focus on images of faces, but our technique is generally applicable. In many areas (such as medicine, astronomy, microscopy, and satellite imagery), sharp, high-resolution images are difficult to obtain due to issues of cost, hardware restriction, or memory limitations . This leads to the capture of blurry, low-resolution images instead. In other cases, images could be old and therefore blurry, or even in a modern context, an image could be out of focus or a person could be in the background. In addition to being visually unappealing, this impairs the use of downstream analysis methods (such as image segmentation, action recognition, or disease diagnosis) which depend on having high-resolution images . In addition, as consumer laptop, phone, and television screen resolution has increased over recent years, popular demand for sharp images and video has surged. This has motivated recent interest in the computer vision task of image super-resolution, the creation of realistic high-resolution (henceforth HR) images that a given low-resolution (LR) input image could correspond to.

While the benefits of methods for image super-resolution are clear, the difference in information content between HR and LR images (especially at high scale factors) hampers efforts to develop such techniques. In particular, LR images inherently possess less high-variance information; details can be blurred to the point of being visually indistinguishable. The problem of recovering the true HR image depicted by an LR input, as opposed to generating a set of potential such HR images, is inherently ill-posed, as the size of the total set of these images grows exponentially with the scale factor . That is to say, many high-resolution images can correspond to the exact same low-resolution image.

Traditional supervised super-resolution algorithms train a model (usually, a convolutional neural network, or CNN) to minimize the pixel-wise mean-squared error (MSE) between the generated super-resolved (SR) images and the corresponding ground-truth HR images . However, this approach has been noted to neglect perceptually relevant details critical to photorealism in HR images, such as texture . Optimizing on an average difference in pixel-space between HR and SR images has a blurring effect, encouraging detailed areas of the SR image to be smoothed out to be, on average, more (pixelwise) correct. In fact, in the case of mean squared error (MSE), the ideal solution is the (weighted) pixel-wise average of the set of realistic images that downscale properly to the LR input (as detailed later). The inevitable result is smoothing in areas of high variance, such as areas of the image with intricate patterns or textures. As a result, MSE should not be used alone as a measure of image quality for super-resolution.

To avoid these issues, we propose a new paradigm for super-resolution. The goal should be to generate realistic images within the set of feasible solutions; that is, to find points which actually lie on the natural image manifold and also downscale correctly. The (weighted) pixel-wise average of possible solutions yielded by the MSE does not generally meet this goal for the reasons previously described. We provide an illustration of this in Figure 2.

Our method generates images using a (pretrained) generative model approximating the distribution of natural images under consideration. For a given input LR image, we traverse the manifold, parameterized by the latent space of the generative model, to find regions that downscale correctly. In doing so, we find examples of realistic images that downscale properly, as shown in 1.

Such an approach also eschews the need for supervised training, being entirely self-supervised with no ‘training’ needed at the time of super-resolution inference (except for the unsupervised generative model). This framework presents multiple substantial benefits. First, it allows the same network to be used on images with differing degradation operators even in the absence of a database of corresponding LR-HR pairs (as no training on such databases takes place). Furthermore, unlike previous methods, it does not require super-resolution task-specific network architectures, which take substantial time on the part of the researcher to develop without providing real insight into the problem; instead, it proceeds alongside the state-of-the-art in generative modeling, with zero retraining needed.

Our approach works with any type of generative model with a differentiable generator, including flow-based models, variational autoencoders (VAEs), and generative adversarial networks (GANs); the particular choice is dictated by the tradeoffs each make in approximating the data manifold. For this work, we elected to use GANs due to recent advances yielding high-resolution, sharp images .

One particular subdomain of image super-resolution deals with the case of face images. This subdomain – known as face hallucination – finds application in consumer photography, photo/video restoration, and more . As such, it has attracted interest as a computer vision task in its own right. Our work focuses on face hallucination, but our methods extend to a more general context.

Because our method always yields a solution that both lies on the natural image manifold and downsamples correctly to the original low-resolution image, we can provide a range of interesting high-resolution possibilities e.g. by making use of the stochasticity inherent in many generative models: our technique can create a set of images, each of which is visually convincing, yet look different from each other, where (without ground truth) any of the images could plausibly have been the source of the low-resolution input.

A new paradigm for image super-resolution. Previous efforts take the traditional, ill-posed perspective of attempting to ‘reconstruct’ an HR image from an LR input, yielding outputs that, in effect, average many possible solutions. This averaging introduces undesirable blurring. We introduce new approach to super-resolution: a super-resolution algorithm should create realistic high-resolution outputs that downscale to the correct LR input.

A novel method for solving the super-resolution task. In line with our new perspective, we propose a new algorithm for super-resolution. Whereas traditional work has at its core aimed to approximate the LR →\rightarrow HR map using supervised learning (especially with neural networks), our approach centers on the use of unsupervised generative models of HR data. Using generative adversarial networks, we explore the latent space to find regions that map to realistic images and downscale correctly. No retraining is required. Our particular implementation, using StyleGAN , allows for the creation of any number of realistic SR samples that correctly map to the LR input.

An original method for latent space search under high-dimensional Gaussian priors. In our task and many others, it is often desirable to find points in a generative model’s latent space that map to realistic outputs. Intuitively, these should resemble samples seen during training. At first, it may seem that traditional log-likelihood regularization by the latent prior would accomplish this, but we observe that the ‘soap bubble’ effect (that much of the density of a high dimensional Gaussian lies close to the surface of a hypersphere) contradicts this. Traditional log-likelihood regularization actually tends to draw latent vectors away from this hypersphere and, instead, towards the origin. We therefore constrain the search space to the surface of that hypersphere, which ensures realistic outputs in higher-dimensional latent spaces; such spaces are otherwise difficult to search.

Related Work

While there is much work on image super-resolution prior to the advent of convolutional neural networks (CNNs), CNN-based approaches have rapidly become state-of-the-art in the area and are closely relevant to our work; we therefore focus on neural network-based approaches here. Generally, these methods use a pipeline where a low-resolution (LR) image, created by down-sampling a high-resolution (HR) image, is fed through a CNN with both convolutional and upsampling layers, generating a super-resolved (SR) output. This output is then used to calculate the loss using the chosen loss function and the original HR image.

Recently, supervised neural networks have come to dominate current work in super-resolution. Dong et al. proposed the first CNN architecture to learn this non-linear LR to HR mapping using pairs of HR-LR images. Several groups have attempted to improve the upsampling step by utilizing sub-pixel convolutions and transposed convolutions . Furthermore, the application of ResNet architectures to super-resolution (started by SRResNet ), has yielded substantial improvement over more traditional convolutional neural network architectures. In particular, the use of residual structures allowed for the training of larger networks. Currently, there exist two general trends: one, towards networks that primarily better optimize pixel-wise average distance between SR and HR, and two, networks that focus on perceptual quality.

2 Loss Functions

Towards these different goals, researchers have designed different loss functions for optimization that yield images closer to the desired objective. Traditionally, the loss function for the image super-resolution task has operated on a per-pixel basis, usually using the L2 norm of the difference between the ground truth and the reconstructed image, as this directly optimizes PSNR (the traditional metric for the super-resolution task). More recently, some researchers have started to use the L1 norm since models trained using L1 loss seem to perform better in PSNR evaluation. The L2 norm (as well as pixel-wise average distances in general) between SR and HR images has been heavily criticized for not correlating well with human-observed image quality . In face super-resolution, the state-of-the-art for such metrics is FSRNet , which used a facial prior to achieve previously unseen PSNR.

Perceptual quality, however, does not necessarily increase with higher PSNR. As such, different methods, and in particular, objective functions, have been developed to increase perceptual quality. In particular, methods that yield high PSNR result in blurring of details. The information required for details is often not present in the LR image and must be ‘imagined’ in. One approach to avoiding the direct use of the standard loss functions was demonstrated in , which draws a prior from the structure of a convolutional network. This method produces similar images to the methods that focus on PSNR, which lack detail, especially in high frequency areas. Because this method cannot leverage learned information about what realistic images look like, it is unable to fill in missing details. Methods that try to learn a map from LR to HR images can try to leverage learned information; however, as mentioned, networks optimized on PSNR are still explicitly penalized for attempting to hallucinate details they are unsure about, thus optimizing on PSNR stills resulting in blurring and lack of detail.

To resolve this issue, some have tried to use generative model-based loss terms to provide these details. Neural networks have lent themselves to application in generative models of various types (especially generative adversarial networks–GANs–from ), to image reconstruction tasks in general, and more recently, to super-resolution. Ledig et al. created the SRGAN architecture for single-image upsampling by leveraging these advances in deep generative models, specifically GANs. Their general methodology was to use the generator to upscale the low-resolution input image, which the discriminator then attempts to distinguish from real HR images, then propagate the loss back to both networks. Essentially, this optimizes a supervised network much like MSE-based methods with an additional loss term corresponding to how fake the discriminator believes the generated images to be. However, this approach is fundamentally limited as it essentially results in an averaging of the MSE-based solution and a GAN-based solution, as we discuss later. In the context of faces, this technique has been incorporated into FSRGAN, resulting in the current perceptual state-of-the-art in face super resolution at ×8\times 8 upscaling factors up to resolutions of 128×128128\times 128. Although these methods use a ‘generator’ and a ‘discriminator’ as found in GANs, they are trained in a completely supervised fashion; they do not use unsupervised generative models.

3 Generative Networks

Our algorithm does not simply use GAN-style training; rather, it uses a truly unsupervised GAN (or, generative model more broadly). It searches the latent space of this generative model for latents that map to images that downscale correctly. The quality of cutting-edge generative models is therefore of interest to us.

As GANs have produced the highest-quality high-resolution images of deep generative models to date, we chose to focus on these for our implementation. Here we provide a brief review of relevant GAN methods with high-resolution outputs. Karras et al. presented some of the first high-resolution outputs of deep generative models in their ProGAN algorithm, which grows both the generator and the discriminator in a progressive fashion. Karras et al. further built upon this idea with StyleGAN, aiming to allow for more control in the image synthesis process relative to the black-box methods that came before it. The input latent code is embedded into an intermediate latent space, which then controls the behavior of the synthesis network with adaptive instance normalization applied at each convolutional layer. This network has 18 layers (2 each for each resolution from 4 ×\times 4 to 1024 ×\times 1024). After every other layer, the resolution is progressively increased by a factor of 2. At each layer, new details are introduced stochastically via Gaussian input to the adaptive instance normalization layers. Without perturbing the discriminator or loss functions, this architecture leads to the option for scale-specific mixing and control over the expression of various high-level attributes and variations in the image (e.g. pose, hair, freckles, etc.). Thus, StyleGAN provides a very rich latent space for expressing different features, especially in relation to faces.

Method

where ∥⋅∥p\|\cdot\|_{p} denotes some lpl^{p} norm.

This is minimized when ISRI_{SR} is an lpl_{p} average of IHRI_{HR} over M∩RM\cap R. In fact, when p=2p=2, this is minimized when

so the optimal ISRI_{SR} is a weighted pixelwise average of the set of high resolution images that downscale properly. As a result, the lack of detail in algorithms that rely only on an lpl_{p} norm cannot be fixed simply by changing the architecture of the network. The problem itself has to be rephrased.

Then we are seeking an image ISR∈M∩RϵI_{SR}\in\mathcal{M}\cap\mathcal{R}_{\epsilon}. The set M∩Rϵ\mathcal{M}\cap\mathcal{R}_{\epsilon} is the set of feasible solutions, because a solution is not feasible if it did not downscale properly and look realistic.

It is also interesting to note that the intersections M∩Rϵ\mathcal{M}\cap\mathcal{R}_{\epsilon} and in particular M∩R0\mathcal{M}\cap\mathcal{R}_{0} are guaranteed to be nonempty, because they must contain the original HR image (i.e., what traditional methods aim to reconstruct).

Central to the problem of super-resolution, unlike general image generation, is the notion of correctness. Traditionally, this has been interpreted to mean how well a particular ground truth image IHRI_{HR} is ‘recovered’ by the application of the super-resolution algorithm SRSR to the low-resolution input ILRI_{LR}, as discussed in the related work section above. This is generally measured by some lpl_{p} norm between ISRI_{SR} and the ground truth, IHRI_{HR}; such algorithms only look somewhat like real images because minimizing this metric drives the solution somewhat nearer to the manifold. However, they have no way to ensure that ISRI_{SR} lies close to M\mathcal{M}. In contrast, in our framework, we never deviate from M\mathcal{M}, so such a metric is not necessary. For us, the critical notion of correctness is how well the generated SR image ISRI_{SR} corresponds to ILRI_{LR}.

We formalize this through the downscaling loss, to explicitly penalize a proposed SR image for deviating from its LR input (similar loss terms have been proposed in ,). This is inspired by the following: for a proposed SR image to represent the same information as a given LR image, it must downscale to this LR image. That is,

where DS(⋅)DS(\cdot) represents the downscaling function.

Our downscaling loss therefore penalizes SRSR the more its outputs violate this,

It is important to note that the downscaling loss can be used in both supervised and unsupervised models for super-resolution; it does not depend on an HR reference image.

2 Latent Space Exploration

How might we find regions of the natural image manifold M\mathcal{M} that map to the correct LR image under the downscaling operator? If we had a differentiable parameterization of the manifold, we could progress along the manifold to these regions by using the downscaling loss to guide our search. In that case, images found would be guaranteed to be high resolution as they came from the HR image manifold, while also being correct as they would downscale to the LR input.

In reality, we do not have such convenient, perfect parameterizations of manifolds. However, we can approximate such a parameterization by using techniques from unsupervised learning. In particular, much of the field of deep generative modeling (e.g. VAEs, flow-based models, and GANs) is concerned with creating models that map from some latent space to a given manifold of interest. By leveraging advances in generative modeling, we can even use pretrained models without the need to train our own network. Some prior work has aimed to find vectors in the latent space of a generative model to accomplish a task; see for creating embeddings and in the context of compressed sensing. (However, as we describe later, this work does not actually search in a way that yields realistic outputs as intended.) In this work, we focus on GANs, as recent work in this area has resulted in the highest quality image-generation among unsupervised models.

Regardless of its architecture, let the generator be called GG, and let the latent space be L\mathcal{L}. Ideally, we could approximate M\mathcal{M} by the image of GG, which would allow us to rephrase the problem above as the following: find a latent vector z∈Lz\in\mathcal{L} with

Experiments

We designed various experiments to assess our method. We focus on the popular problem of face hallucination, enhanced by recent advances in GANs applied to face generation. In particular, we use Karras et al.’s pretrained Face StyleGAN (trained on the Flickr Face HQ Dataset, or FFHQ) . We adapted the implementation found at in order to transfer the original StyleGAN-FFHQ weights and model from TensorFlow to PyTorch . For each experiment, we used 100100 steps of spherical gradient descent with a learning rate of 0.40.4 starting with a random initialization. Each image was therefore generated in ∼5{\sim}5 seconds on a single NVIDIA V100 GPU.

We evaluated our procedure on the well-known high-resolution face dataset CelebA HQ. (Note: this is not to be confused with CelebA, which is of substantially lower resolution.) We performed these experiments using scale factors of 64×64\times, 32×32\times, and 8×8\times. For our qualitative comparisons, we upscale at scale factors of both 8×8\times and 64×64\times, i.e., from 16×1616\times 16 to 128×128128\times 128 resolution images and 1024×10241024\times 1024 resolution images. The state-of-the-art for face super-resolution in the literature prior to this point was limited to a maximum of 8×8\times upscaling to a resolution of 128×128128\times 128, thus making it impossible to directly make quantitative comparisons at high resolutions and scale factors. We followed the traditional approach of training the supervised methods on CelebA HQ. We tried comparing with supervised methods trained on FFHQ, but they failed to generalize and yielded very blurry and distorted results when evaluated on CelebA HQ; therefore, in order to compare our method with the best existing methods, we elected to train the supervised models on CelebA HQ instead of FFHQ.

2 Qualitative Image Results

Figure 5 shows qualitative results to demonstrate the visual quality of the images from our method. We observe levels of detail that far surpass competing methods, as exemplified by certain high frequency regions (features like eyes or lips). More examples and full-resolution images are in the appendix.

3 Quantitative Comparison

Here we present a quantitative comparison with state-of-the-art face super-resolution methods. Due to constraints on the peak resolution that previous methods can handle, evaluation methods were limited, as detailed below.

We conducted a mean-opinion-score (MOS) test as is common in the perceptual super-resolution literature . For this, we had 40 raters examine images upscaled by 6 different methods (nearest-neighbors, bicubic, FSRNet, FSRGAN, and our PULSE). For this comparison, we used a scale factor of 88 and a maximum resolution of 128×128128\times 128, despite our method’s ability to go substantially higher, due to this being the maximum limit for the competing methods. After being exposed to 20 examples of a 1 (worst) rating exemplified by nearest-neighbors upsampling, and a 5 (best) rating exemplified by high-quality HR images, raters provided a score from 1-5 for each of the 240 images. All images fell within the appropriate ϵ=1e−3\epsilon=1e-3 for the downscaling loss. The results are displayed in Table 1.

PULSE outperformed the other methods and its score approached that of the HR dataset. Note that the HR’s 3.74 average image quality reflects the fact that some of the HR images in the dataset had noticeable artifacts. All pairwise differences were highly statistically significant (p<10−5p<10^{-5} for all 15 comparisons) by the Mann-Whitney-U test. The results demonstrate that PULSE outperforms current methods in generating perceptually convincing images that downscale correctly.

To provide another measure of perceptual quality, we evaluated the Naturalness Image Quality Evaluator (NIQE) score , previously used in perceptual super-resolution . This no-reference metric extracts features from images and uses them to compute a perceptual index (lower is better). As such, however, it only yields meaningful results at higher resolutions. This precluded direct comparison with FSRNet and FSRGAN, which produce images of at most 128×128128\times 128 pixels.

We evaluated NIQE scores for each method at a resolution of 1024×10241024\times 1024 from an input resolution of 16×1616\times 16, for a scale factor of 6464. All images for each method fell within the appropriate ϵ=1e−3\epsilon=1e-3 for the downscaling loss. The results are in Table 2.

PULSE surpasses even the CelebA HQ images in terms of NIQE here, further showing the perceptual quality of PULSE’s generated images. This is possible as NIQE is a no-reference metric which solely considers perceptual quality; unlike reference metrics like PSNR, performance is not bounded above by that of the HR images typically used as reference.

4 Image Sampling

As referenced earlier, we initialize the point we start at in the latent space by picking a random point on the sphere. We found that we did not encounter any issues with convergence from random initializations. In fact, this provided us one method of creating many different outputs with high-level feature differences: starting with different initializations. An example of the variation in outputs yielded by this process can be observed in Figure 3.

Furthermore, by utilizing a generative model with inherent stochasticity, we found we could sample faces with fine-level variation that downscale correctly; this procedure can be repeated indefinitely. In our implementation, we accomplish this by resampling the noise inputs that StyleGAN uses to fill in details within the image.

Robustness

The main aim of our algorithm is to perform perceptually realistic super-resolution with a known downscaling operator. However, we find that even for a variety of unknown downscaling operators, we can apply our method using bicubic downscaling as a stand-in for more substantial degradations applied–see Figure 6. In this case, we provide only the degraded low-resolution image as input. We find that the output downscales approximately to the true, non-noisy LR image (that is, the bicubically downscaled HR) rather than to the degraded LR given as input. This is desired behavior, as we would not want to create an image that matches the additional degradations. PULSE thus implicitly denoises images. This is due to the fact that we restrict the outputs to only realistic faces, which in turn can only downscale to reasonable LR faces. Traditional supervised networks, on the other hand, are sensitive to added noise and changes in the domain and must therefore be explicitly trained with the noisy inputs (e.g., ).

Bias

While we initially chose to demonstrate PULSE using StyleGAN (trained on FFHQ) as the generative model for its impressive image quality, we noticed some bias when evaluated on natural images of faces outside of our test set. In particular, we believe that PULSE may illuminate some biases inherent in StyleGAN. We document this in a more structured way with a model card in Figure PULSE: Self-Supervised Photo Upsampling via Latent Space Exploration of Generative Models, where we also examine success/failure rates of PULSE across different subpopulations. We propose a few possible sources for this bias:

Bias inherited from latent space constraint: If StyleGAN pathologically placed people of color in areas of lower density in the latent space, bias would be introduced by PULSE’s constraint on the latent space which is necessary to consistently generate high resolution images. To evaluate this, we ran PULSE with different radii for the hypersphere PULSE searches on, corresponding to different samples. This did not seem to have an effect.

Failure to converge: In the initial code we released on GitHub, PULSE failed to return “no image found” when at the end of optimization it still did not find an image that downscaled correctly (within ϵ\epsilon). The concern could therefore be that it is harder to find images in the outputs of StyleGAN that downscale to people of color than to white people. To test this, we found a new dataset with better representation to evaluate success/failure rates on, “FairFace: Face Attribute Dataset for Balanced Race, Gender, and Age” . This dataset was labeled by third-party annotators on these fields. We sample 100 examples per subgroup and calculate the success/failure rate across groups with ×64\times 64 downscaling after running PULSE 5 times per image. The results of this experiment can be found in Table 3. There is some variation in these percentages, but it does not seem to be the primary cause of what we observed. Note that this metric is lacking in that it only reports whether an image was found - which does not reflect the diversity of images found over many runs on many images, an important measure that was difficult to quantify.

Bias inherited from optimization: This would imply that the constrained latent space contains a wide range of images of people of color but that PULSE’s optimization procedure does not find them. However, if this is the case then we should be able to find such images with enough random initializations in the constrained latent space. We ran this experiment and this also did not seem to have an effect.

Bias inherited from StyleGAN: Some have noted that it seems more diverse images can be found in an augmented latent space of StyleGAN per . However, this is not close to the set of images StyleGAN itself generates when trained on faces: for example, in the same paper, the authors display images of unrelated domains (such as cats) being embedded successfully as well. In our work, PULSE is constrained to images StyleGAN considers realistic face images (actually, a slight expansion of this set; see below and Appendix for an explanation of this).

More technically: in StyleGAN, they sample a latent vector zz, which is fed through the mapping network to become a vector ww, which is duplicated 18 times to be fed through the synthesis network. In , they instead find 18 different vectors to feed through the synthesis network that correspond to any image of interest (whether faces or otherwise). In addition, while each of these latent vectors would have norm ≈512\approx\sqrt{512} when sampled, this augmented latent space allows them to vary freely, potentially finding points in the latent space very far from what would have been seen in training. Using this augmented latent space therefore removes any guarantee that the latent recovered corresponds to a realistic image of a face.

In our work, instead of duplicating the vector 18 times as StyleGAN does, we relax this constraint and encourage these 18 vectors to be approximately equal to each other so as to still generate realistic outputs (see Appendix). This relaxation means the set of images PULSE can generate should be broader than the set of images StyleGAN could produce naturally. We found that loosening any of StyleGAN’s constraints on the latent space further generally led to unrealistic faces or images that were not faces.

Overall, it seems that sampling from StyleGAN yields white faces much more frequently than faces of people of color, indicating more of the prior density may be dedicated to white faces. Recent work by Salminen et al. describes the implicit biases of StyleGAN in more detail, which seem to confirm these observations. In particular, we note their analysis of the demographic bias of the outputs of the model:

“Results indicate a racial bias among the generated pictures, with close to three-[fourths] (72.6%) of the pictures representing White people. Asian (13.8%) and Black (10.1%) are considerably less frequent, while Indians represent only a minor fraction of the pictures (3.4%).”

This bias extends to any downstream application of StyleGAN, including the implementation of PULSE using StyleGAN.

Discussion and Future Work

Through these experiments, we find that PULSE produces perceptually superior images that also downscale correctly. PULSE accomplishes this at resolutions previously unseen in the literature. All of this is done with unsupervised methods, removing the need for training on paired datasets of LR-HR images. The visual quality of our images as well as MOS and NIQE scores demonstrate that our proposed formulation of the super-resolution problem corresponds with human intuition. Starting with a pre-trained GAN, our method operates only at test time, generating each image in about 5 seconds on a single GPU. However, we also note significant limitations when evaluated on natural images past the standard benchmark.

One reasonable concern when searching the output space of GANs for images that downscale properly is that while GANs generate sharp images, they need not cover the whole distribution as, e.g., flow-based models must. In our experiments using CelebA and StyleGAN, we did not observe any manifestations of this, which may be attributable to bias – see Section 6 (The “mode collapse” behavior of GANs may exacerbate dataset bias and contribute to the results described in Section 6 and the model card, Figure PULSE: Self-Supervised Photo Upsampling via Latent Space Exploration of Generative Models.) Advances in generative modeling will allow for generative models with better coverage of larger distributions, which can be directly used with PULSE without modification.

Another potential concern that may arise when considering this unsupervised approach is the case of an unknown downscaling function. In this work, we focused on the most prominent SR use case: on bicubically downscaled images. In fact, in many use cases, the downscaling function is either known analytically (e.g., bicubic) or is a (known) function of hardware. However, methods have shown that the degradations can be estimated in entirely unsupervised fashions for arbitrary LR images (that is, not necessarily those which have been downscaled bicubically) . Through such methods, we can retain the algorithm’s lack of supervision; integrating these is an interesting topic for future work.

Conclusions

We have established a novel methodology for image super-resolution as well as a new problem formulation. This opens up a new avenue for super-resolution methods along different tracks than traditional, supervised work with CNNs. The approach is not limited to a particular degradation operator seen during training, and it always maintains high perceptual quality.

Acknowledgments: Funding was provided by the Lord Foundation of North Carolina and the Duke Department of Computer Science. Thank you to the Google Cloud Platform research credits program.

References