Meta-Transfer Learning for Zero-Shot Super-Resolution
Jae Woong Soh, Sunwoo Cho, Nam Ik Cho
Introduction
SISR, which is to find a plausible HR image from its counterpart LR image, is a long-standing problem in low-level vision area. Recently, the remarkable success of CNNs brought attention to the research community, and hence numerous CNN-based SISR methods have exhibited large performance leap VDSR; SRGAN; EDSR; RDN; CARN; RCAN; NatSR; SRFBN; DBPN; OISR. Most of the recent state-of-the-art (SotA) CNN-based methods are based on a large number of external training dataset and self-supervised settings with known degradation model, e.g., “bicubic” downsampling. Impressively, the recent SotA CNNs show significant PSNR gains compared to the conventional large size of models for the noise-free “bicubic” downsampling condition. However, in real-world situations, when the LR image has distant statistics in downsampling kernels and noises, the recent methods produce undesirable artifacts and show inferior results due to the domain gap. Moreover, their number of parameters and memory overheads are usually too large to be used in real applications.
Besides, non-local self-similarity in scale and across multi-scale, which is the internal recurrence of information within a single image, is one of the strong natural image priors. Therefore it has long been used in image restoration tasks, including image denoising NLM; BM3D and super-resolution Non-Param; Self-Exemplar. Additionally, the powerful image prior of non-local property is embedded into network architecture Non-local; NLRN; RNAN by implicitly learning such priors to boost the performance of the networks further. Also, some works to learn internal distribution have been proposed ZSSR; SinGAN; InGAN. Moreover, there have been many studies to combine the advantages of external and internal information for image restoration comb1; comb2; comb3; comb4.
Recently, ZSSR ZSSR has been proposed for zero-shot super-resolution, which is based on the zero-shot setting to exploit the power of CNN but can be easily adapted to the test image condition. Interestingly, ZSSR learns the internal non-local structure of the test image, i.e., deep internal learning. Thus it outperforms external-based CNNs in some regions where the recurrences are salient. Also, ZSSR is highly flexible that it can address any blur kernels, and thus easily adapted to the conditions of test images.
However, ZSSR has a few limitations. First, it requires thousands of backpropagation gradient updates at test time, which requires considerable time to get the result. Also, it cannot fully exploit the large-scale external dataset, and rather it depends only on internal structure and patterns, which lacks in the number of total examples. Eventually, this leads to inferior results in most of the regions with general patterns compared to the external-based methods.
On the other hand, meta-learning or learning to learn fast has recently attracted many researchers. Meta-learning aims to address a problem that artificial intelligence is hard to learn new concepts quickly with a few examples, unlike human intelligence. In this respect, meta-learning is jointly merged with few-shot learning, and many methods with this approach have been proposed Prototypical; Matching; Relation; TADAM; SNAIL; MAML; GRAD1; GRAD2; MTL. Among them, Model-Agnostic Meta-Learning (MAML) MAML has shown great impact, showing SotA performance by learning the optimal initial state of the model such that the base-learner can fast adapt to a new task within a few gradient steps. MAML employs the gradient update as meta-learner, and the same author analyzed that gradient descent can approximate any learning algorithm Meta_Grad. Moreover, Sun et al. MTL have jointly utilized MAML with transfer learning to exploit large-scale data for few-shot learning.
Inspired by the above-stated works and ZSSR, we present Meta-Transfer Learning for Zero-Shot Super-Resolution (MZSR), which is kernel-agnostic. We found that simply employing transfer learning or fine-tuning from a pre-trained network does not yield plausible results. As ZSSR only has a meta-test step, we additionally adopt a meta-training step to make the model adapt fast to new blur kernel scenarios. Additionally, we adopt transfer learning in advance to fully utilize external samples, further leveraging the performance. In particular, transfer learning with the help of a large-scale synthetic dataset (“bicubic” degradation setting) is first performed for the external learning of natural image priors. Then, meta-learning plays a role in learning task-level knowledge with different downsampling kernels as different tasks. At the meta-test step, simple self-supervised learning is conducted to learn image-specific information within a few gradient steps. As a result, we can exploit both external and internal information. Also, by leveraging the advantages of ZSSR, we may use a lightweight network, which is flexible to different degradation conditions of LR images. Furthermore, our method is much faster than ZSSR, i.e., it quickly adapts to new tasks within a few gradient steps, while ZSSR requires thousands of updates.
In summary, our overall contribution is three-fold:
We present a novel training scheme based on meta-transfer learning, which learns an effective initial weight for fast adaptation to new tasks with the zero-shot unsupervised setting.
By using external and internal samples, it is possible to leverage the advantages of both internal and external learning.
Our method is fast, flexible, lightweight and unsupervised at meta-test time, hence, eventually can be applied to real-world scenarios.
Related Work
SISR is based on the image degradation model as
where , , , , , and denote HR, LR image, blur kernel, convolution, decimation with scaling factor of , and white Gaussian noise, respectively. It is notable that diverse degraded conditions can be found in real-world scenes, with various unknown , , and .
Recently, numerous CNN-based networks have been proposed to super-resolve LR image with known downsampling kernel VDSR; SRGAN; EDSR; DBPN; RDN; CARN; NatSR; SRFBN; OISR. They show extreme performances in “bicubic” downsampling scenarios but suffer in non-bicubic cases due to the domain gap. To cope with multiple degradation kernels, SRMD SRMD has been proposed. With additional inputs of kernel and noise information, SRMD outperforms other SISR methods in non-bicubic conditions. Also, IKC BlindSR has been proposed for blind super-resolution. On the other hand, ZSSR ZSSR has been proposed to learn image specific internal structure with CNN, and has shown that it can be applied to real-world scenes due to its flexibility.
2 Meta-Learning
In recent years, diverse meta-learning algorithms have been proposed. They can be categorized into three groups. The first group is metric based methods Prototypical; Relation; Matching, which is to learn metric space in which learning is efficient within a few samples. The second group is memory network-based methods memory; TADAM; SNAIL, where the network learns across task knowledges and well generalizes to unseen tasks. The last group is optimization based methods, where gradient descent plays a role as a meta-learner optimization GRAD1; GRAD2; Meta_Grad; MAML. Among them, MAML MAML has shown a great impact on the research community, and several variants have been proposed Reptile; MTL; MAML++; MAML_Embed. MAML inherently requires second-order derivative terms, and the first-order algorithm has also been proposed in Reptile. Also, to cope with the instability of MAML training, MAML++ MAML++ has been proposed. Moreover, MAML within embedded space has been proposed MAML_Embed. In this paper, we employ MAML scheme for fast adaptation of zero-shot super-resolution.
Preliminary
We introduce self-supervised zero-shot super-resolution and meta-learning schemes with notations, following related works ZSSR; MAML.
ZSSR ZSSR is totally unsupervised or self-supervised. Two phases of training and test are both held in runtime. In training phase, the test image is downsampled with desired kernel to generate “LR son” denoted as , and becomes the HR supervision, “HR father.” Then, the CNN is trained with the LR-HR pairs generated by a single image. The training solely depends on the test image, thus learns specific internal information to given image statistics. In the test phase, the trained CNN then works as a feedforward network, and the test input image is fed to the CNN to get the super-resolved image .
Meta-Learning
Meta-learning has two phases: meta-training and meta-test. We consider a model , which is parameterized by , that maps inputs to outputs . The goal of meta-training is to make the model to be able to adapt to a large number of different tasks. A task is sampled from a task distribution for meta-training. Within a task, training samples are used to optimize the base-learner with a task-specific loss and test samples are used to optimize the meta-learner. In meta-test phase, the model quickly adapts to a new task with the help of meta-learner. MAML MAML employs a simple gradient descent algorithm as the meta-learner and seeks to find an initial transferable point where a few gradient updates lead to a fast adaptation of the model to a new task.
In our case, the input and the output are and . Also, diverse blur kernels constitute the task distribution, where each task corresponds to the super-resolution of an image degraded by a specific blur kernel.
Method
The overall scheme of our proposed MZSR is shown in Figure 2. As shown, our method consists of three steps: large-scale training, meta-transfer learning, and meta-test.
This step is similar to the large-scale ImageNet ImageNet pre-training for object recognition. In our case, we adopt DIV2K DIV2K which is a high-quality dataset . Using known “bicubic” degradation, we first synthesized large number of paired dataset , denoted as . Then, we trained the network to learn super-resolution of “bicubic” degradation model by minimizing the loss,
which is the pixel-wise L1 loss EDSR; ZSSR between prediction and the ground-truth.
The large-scale training has contributions within two respects. First, as super-resolution tasks share similar properties, it is possible to learn efficient representations that implicitly represent natural image priors of high-resolution images, thus making the network ease to be learned. Second, as MAML MAML is known to show some unstable training, we ease the training phase of meta-learning with the help of well pre-trained feature representations.
2 Meta-Transfer Learning
Since ZSSR is trained with the gradient descent algorithm, it is possible to introduce an optimization-based meta-training step with the help of gradient descent algorithm, which is proven to be a universal learning algorithm Meta_Grad.
In this step, we seek to find a sensitive and transferable initial point of the parameter space where a few gradient updates lead to large performance improvements. Inspired by MAML, our algorithm mostly follows MAML but with several modifications.
Unlike MAML, we adopt different settings for meta-training and meta-test. In particular, we use the external dataset for meta-training, whereas internal learning is adopted for meta-test. This is because we intend our meta-learner to more focus on the kernel-agnostic property with the help of a large-scale external dataset.
We synthesize dataset for meta-transfer learning, denoted as . consists of pairs, , with diverse kernel settings. Specifically, we used isotropic and anisotropic Gaussian kernels for the blur kernels. We consider a kernel distribution , where each kernel is determined by a covariance matrix . it is chosen to have a random angle , and two random eigenvalues , where denotes the scaling factor. Precisely, the covariance matrix is expressed as
Eventually, we train our meta-learner based on . We may divide into two groups: for task-level training, and for task-level test.
In our method, adaptation to a new task with respect to the parameters is one or more gradient descent updates. For one gradient update, new adapted parameters is then
where is the task-level learning rate. The model parameters are optimized to achieve minimal test error of with respect to . Concretely, the meta-objective is
Meta-transfer optimization is performed using Eq. 6, which is to learn the knowledge across task. Any gradient-based optimization can be used for meta-transfer training. For stochastic gradient descents, the parameter update rule is expressed as
3 Meta-Test
The meta-test step is exactly the zero-shot super-resolution. As evidence in ZSSR, this step enables our model to learn internal information within a single image. With a given LR image, we downsample it with corresponding downsampling kernel (kernel estimation algorithms Non-Param; KernelEst can be adopted for blind scenario) to generate and perform a few gradient updates with respect to the model parameter using a single pair of “LR son” and a given image. Then, we feed a given LR image to the model to get a super-resolved image.
4 Algorithm
Algorithm 1 demonstrates the process of our meta-transfer training procedures of Section 4.1 and 4.2. Lines 3-7 is the large-scale training stage. Lines 11-14 is the inner loop of meta-transfer learning where the base-learner is updated to task-specific loss. Lines 15-16 presents the meta-learner optimization.
Algorithm 2 presents the meta-test step, which is the zero-shot super-resolution. A few gradient updates () are performed while meta-test, and the super-resolved image is obtained with final updated parameters.
Experiments
For the CNN, we adopt a simple -layer CNN architecture with residual learning following ZSSR ZSSR. Its number of parameters is K. For meta-transfer training, we use DIV2K DIV2K for the high-quality dataset and we set and for entire training. For the inner loop, we conducted gradient updates, i.e. unrolling steps, to obtain adapted parameters. We extracted training patches with a size of . To cope with gradient vanishing or exploding problems due to the unrolling process of base learners, we utilize the weighted sum of losses from each step, i.e., providing supervision of additional losses to each unrolling step MAML++. At the initial point, we evenly weigh the losses and decayed the weights except for the last unrolling step. In the end, the weighted loss converges to our final training task loss. We employ ADAM ADAM optimizer as our meta-optimizer. As the subsampling process () can be the direct method ZSSR or the bicubic subsampling SRMD; BlindSR, we trained two models for different subsampling methods: direct and bicubic.
2 Evaluations on “Bicubic” Downsampling
We evaluate our method with several recent SotA SISR methods, including supervised and unsupervised methods on famous benchmarks: Set5 Set5, BSD100 B100, and Urban100 Self-Exemplar. We measure PSNR and SSIM SSIM in Y-channel of YCbCr colorspace.
The overall results are shown in Table 1. CARN CARN and RCAN RCAN, which are trained for “bicubic” downsampling condition, show extremely overwhelming performances. Since the training scenario and the test scenario exactly match each other, supervision on external samples could boost the performance of CNN. On the other hands, ZSSR ZSSR and our methods show improvements against bicubic interpolation but not as good as the supervised ones, because both methods are trained within the unsupervised or self-supervised regime. Our methods show comparable results to ZSSR within only one single gradient descent update.
3 Evaluations on Various Blur Kernels
In this section, we demonstrate the results on various blur kernel conditions. We assume four scenarios: severe aliasing, isotropic Gaussian, unisotropic Gaussian, and isotropic Gaussisan followed by bicubic subsampling. Precisely, the methods are
: isotropic Gaussian blur kernel with width followed by direct subsampling.
: isotropic Gaussian blur kernel with width followed by direct subsampling.
: anisotropic Gaussian with widths and with from Eq. 3, followed by direct subsampling.
: isotropic Gaussian blur kernel with width followed by bicubic subsampling.
The results are shown in Table 2. As the SotA method RCAN RCAN is trained on “bicubic” scenario, it shows inferior performance due to domain discrepancy and lack of flexibility.
For the case of aliasing (), RCAN results are even worse than a simple bicubic interpolation method due to inconsistency between training and test condition. IKC We reimplemented the code and retrained with DIV2K dataset. BlindSR is trained for bicubic subsampling, it never sees aliased images during training. Thus, it also shows a severe performance drop. On the other hand, ZSSR We used the official code but without gradual configuration. ZSSR shows quite improved results due to its flexibility. However, it requires thousands of gradient updates, which require a large amount of time. Also, it starts from a random initial point and thus does not guarantee the same results for multiple tests. As shown in Table 2, our methods are comparable to others even with one single gradient update. Interestingly, our MZSR never sees the kernel with , but the CNN quickly adapts to specific image condition. In other words, compared to other methods, our method is more robust to extrapolation.
For other cases, which are isotropic and anisotropic Gaussian, our methods outperform others with a significantly large gap. In these cases, other methods have performance gains compared to bicubic interpolation, but the differences are minor. Similar tendencies of aliasing cases can be found in all other scenarios. Interestingly, RCAN RCAN shows slightly improved results compared to bicubic interpolation. Also, as the condition between training and test is consistent, IKC BlindSR shows comparable results. Our methods also show remarkable performance in the case of bicubic subsampling condition. From the extensive experimental results, we believe that our MZSR is a fast, flexible, and accurate method for super-resolution.
4 Real Image Super-Resolution
To show the effectiveness of the proposed MZSR, we also conduct experiments on real images. Since there are no ground-truth images for real images, we only present the visual comparisons. Due to the page limit, all the comparisons on real images are presented in supplementary material.
Discussion
For ablation investigation, we train several models with different configurations. We assess the average PSNR results on Set5, which are shown in Figure 3. Interestingly, the initial point of our method shows the worst performance, but in one iteration, our method quickly adapts to the image condition and shows the best performance among the compared methods. Other methods sometimes show a slow increase in performance. In other words, they are not as flexible as ours in adapting to new image conditions.
We visualized the result at the initial point and after one gradient update in Figure 4. As shown, the result of the initial point of MZSR is weird, but within one iteration, it is highly improved. On the other hand, the result of a pre-trained network is more natural than MZSR, but its improvement after one gradient update is minor. Furthermore, it is shown that the performance of our method increases as the gradient descent update progresses, despite the fact that it is trained for maximum performance after five gradient steps. This result suggests that with more gradient update iterations, we might expect more of the performance improvements.
2 Multi-scale Models
We additionally trained a multi-scale model with the scaling factors . The results on show worse results comparable to single-scale model as shown in Table 3. With multiple scaling factors, the task distribution becomes more complex, in which the meta-learner struggles to capture such regions that are suitable for fast adaptation.
Moreover, when meta-testing larger scaling factors, the size of becomes too small to provide enough information to the CNN. Hence, the CNN rarely utilizes information from a very small LR son image. Importantly, as our CNN learns internal information of CNN, such images with multi-scale recurrent patterns show plausible results even with large scaling factors, as shown in Figure 5.
3 Complexity
We evaluate the overall model and time complexities for several comparisons, and the results are shown in Table 4. We measure time on the environment of NVIDIA Titan XP GPU. Two fully-supervised feedforward networks for “bicubic” degradation, CARN and RCAN, require a large number of parameters. Even though CARN is proposed as a lightweight network which requires one-tenth of parameters compared to RCAN, it still requires much more parameters compared to unsupervised networks. However, the time consumptions for both model are quite comparable, because only feedforward computation is involved.
On the other hand, ZSSR, which is totally unsupervised, requires much less number of parameters due to the image-specific CNN. However, it requires thousands of forward and backward pass to get a super-resolved image, i.e., a large amount of time exceeding a practical extent. Our method MZSR with a single gradient update requires the shortest time among comparisons. Also, even with iterations of the backward pass, our method still shows comparable time consumption against CARN.
Conclusion
In this paper, we have presented a fast, flexible, and lightweight self-supervised super-resolution method by exploiting both external and internal samples. Specifically, we adopt an optimization-based meta-learning method jointly with transfer learning to seek an initial point that is sensitive to different conditions of blur kernels. Therefore, our method can quickly adapt to specific image conditions within a few gradient updates. From our extensive experiments, we show that our MZSR outperforms other methods, including ZSSR, which requires thousands of gradient descent iterations. Furthermore, we demonstrate the effectiveness of our method with complexity evaluation. Yet, there are lots of parts that can be improved from our work such as network architecture, learning strategies, and multi-scale model, and we leave these as future works. Our code is publicly available at https://www.github.com/JWSoh/MZSR.
This research was supported in part by Projects for Research and Development of Police science and Technology under Center for Research and Development of Police science and Technology and Korean National Police Agency (PA-C000001), and in part by Samsung Electronics Co., Ltd.
References
Appendix
A. Evaluation on Scaling Factor ×4\times 4
To evaluate the performance on large scaling factors, we demonstrate the results on scaling factor with isotropic Gaussian kernel with width in Table 5. As shown, our methods show comparable results to others even with one gradient update, for large scaling factors too. Also, we found that multi-scale model shows worse results than a single-scale model as evidenced in the scaling factor .
B. Effects of Kernels on Meta-test Time
To evaluate the effects of input kernels on meta-test time, we obtained several results by feeding various kernels. The results are shown in Figure 6. It is obvious that kernel mismatch degrades the output result severely. Especially, when the input kernel largely deviates from the true kernel, the result is not very pleasing as shown in Figure 6(a) and (b). However, if the input kernel has similar shape as the true kernel then the result looks quite plausible as shown in Figure 6(c). In conclusion, the kernel estimation or knowing the true kernel is crucial for the performance gain with our method.
C. Visualization
To show the effectiveness of our MZSR, we visualize some results including scenarios with synthetic blur kernels and real-world images. Figure 7 and 8 are the results on synthetic blur kernels. 9 is the result on a real-world image.