AdderSR: Towards Energy Efficient Image Super-Resolution
Dehua Song, Yunhe Wang, Hanting Chen, Chang Xu, Chunjing Xu, Dacheng Tao
Introduction
Single image super-resolution (SISR) is a typical computer vision task which aims at reconstructing a high-resolution (HR) image from a low-resolution (LR) image. SISR is a very popular image signal processing task in real-world applications such as smart phones and mobile cameras. Due to the hardware constrains of these portable devices, it is necessary to develop SISR models with low computation cost and high visual quality.
Recently, deep convolutional neural network (CNN) has dramatically boosted the performance of SISR. The first super-resolution convolutional neural network (SRCNN) contains only three convolutional layers with about 57K parameters. Then, the capacity of DCNN was amplified with the increasing of depth and width (channel number), resulting in notable improvement of super-resolution. The parameters and computation cost of recent DCNN are increased accordingly. For example, the residual dense network (RDN) contains 22M parameters and requires about 10,192G FLOPs (floating-number operations) for processing only one image. Compared with neural networks for visual recognition (\eg, ResNet-50 ), models for SISR have much higher computational complexities due to the larger feature map sizes. These massive calculations will consume much energy and reduce the enduration time of mobile devices.
In order to address the aforementioned problem, a series of approaches have been proposed to compress and accelerate deep convolutional neural networks. The prominent compression methods such as filter pruning and knowledge distillation reduce the computation by narrowing or shallowing the network. On the other hand, quantization methods devote to reducing the computation complexity of multiplications while preserving the architecture of the original neural network. Wherein, binarization is a specific case that weights and activations in networks are represented as , which can significantly reduce the energy and memory consumptions. However, the binarized network often cannot maintain the accuracy of full precision network, especially for super-resolution task . Recently, Chen et al. proposed a novel AdderNet which replaces the multiplication operations by additions. Since the complexity of additions is much lower than that of multiplications, this work motivates us to utilize AdderNet for constructing energy efficient SISR models.
To maximally excavate the potential for exploiting AdderNets to establish SISR models, we first analyze the theoretical difficulties for applying the additions into SISR tasks. Specifically, input and output features in any two neighbor layers in SISR models are very close with the similar global texture and color information as shown in Figure 2. However, the identify mapping cannot be learned by a one-layer adder network. Thus, we suggest to insert self-shortcuts and formulate new adder models for the SISR task. Moreover, we find that the high-pass filter is also hard to approximate by adder units. We then develop a learnable power activation. By exploiting these two techniques, we replace the conventional convolution filters in modern SISR networks by adder filters and establish AdderSR models accordingly. The effectiveness of the proposed SISR networks using additions is verified on several benchmark datasets. We can obtain comparable performance (\ie, PSNR values and visual quality) using AdderSR models with that of the CNN baselines. Meanwhile, we can reduce more than of the overall energy consumptions of these neural networks. Fig. 1 indicates the superior performance of AdderSR networks with moderate energy.
The rest of this paper is organized as follows. We briefly investigate the related works on neural network compression and energy-efficient approaches in Section 2. In Section 3, we present the motivation of using additions in SISR and establish AdderSR models. Section 4 illustrates both the quantitative and qualitative results on benchmarks and Section 5 concludes the paper.
Related Works
SISR is an important computer vision task, which has broad applications such as photo capture, surveillance and entertainment. In the last decades, numerous prominent approaches have been proposed to solve this ill-posed problem and obtained a tremendous progress in the performance. Unfortunately, current SISR methods require massive computations to process an input image. The energy consumption has become a thorny issue restricting the application of SISR on mobile devices.
Model compression has been investigated for many years and vast novel methods have been proposed. These methods can be roughly divided into four categories: network pruning, efficient filter design, neural architecture search (NAS) and knowledge distillation. Pruning aims at reducing the redundancy of filters so as to decrease the computation of the original model. Hou et al. proposed a new pruning criterion for SISR which judged redundant channels with the discriminant information. The most conventional method is to design efficient block (\eg GhostNet , CARN , IDN , MAFFSRN ) with efficient operators (\eggroup convolution, convolution). Furthermore, NAS has been employed to exploit efficient SR neural architecture automatically. In addition, knowledge distillation can transfer the information from large models to improve the performance of the tiny model. Gao et al. proposed a novel knowledge distillation scheme for SISR and boosted the performance.
2 Low-Cost Computation
In addition to reduce the computation of models, there are two kinds of approaches that decrease the energy consumption while preserving the architecture of original network. Model quantization saves energy by reducing the number of bits required to represent each weight or feature component. Wherein, binarization is a specific case that weights and activations in networks are represented as . Li et al. proposed a novel quantization scheme for SISR to acquire a large dynamic quantization range. Xin et al. designed a bit-accumulation mechanism to alleviate the quantization error of binary SISR networks. Unfortunately, the quantized network often cannot maintain the accuracy of super-resolution networks. On the other hand, many research works have indicated that add addition consumes fewer energy than multiplication. Recently, Chen et al. pioneered a novel method to reduce the power dissipation of network by taking the place of the multiplication with add operation. It achieved marginal loss of accuracy on classification tasks without any multiplication in convolutional layers. Then, kernel based progressive distillation promoted the accuracy of AdderNets even superior to that of standard CNNs. It is attractive to construct energy efficient SISR models with AdderNet.
AdderNet for Image Super-Resolution
AdderNet reduces the energy consumption of classification networks significantly while achieving comparable performance. We aim to inherit this huge success to the image super-resolution task, which often has higher energy consumption and computational complexity. In SISR, there are two important properties that should be ensured by adder neural networks: the similarity between the input and output features of each convolutional layer, and the enhancement of details \wrthigh frequency information. These two properties are illustrated in Fig. 2.
Here we firstly briefly introduce the single image super-resolution tasks using deep learning methods and then discuss the difficulty for directly using AdderNets to construct energy efficient SR models.
SISR aims at reconstructing a high-resolution image from a low-resolution image. It is a typical ill-posed reverse problem since an infinite number of high-resolution images could generate the same low-resolution image with downsampling. The objective function for the conventional SISR task can be formulated as:
where is the observation data, \ie, the low-resolution image, and is the desired high-resolution image, denotes the used priori such as smooth and additive noise, is the tradeoff parameter.
Recently, Dong et al. first introduced deep learning method to super-resolution and achieved much better performance than traditional methods. Along with the improvement of super-resolution performance, the parameters and computations of SR networks grew rapidly, which seriously limited the efficiency of model executing on mobile devices. Hence, another research direction is to deploy efficient SR networks. Quantization, knowledge distillation, efficient operator designing and NAS have been explored to exploit efficient and accurate SR model. However, energy consumptions required by these portable SISR models are still much expensive for real-world mobile devices.
where is the absolute value function. and denote the spatial location of features. is the index of output channels. Although Eq. 2 shows comparable performance on image classification tasks, the SISR problem defined in Eq. 1 is quite different to the conventional recognition task. For example, we need to ensure that the output results maintain the original texture in , which cannot be easily learned by Eq. 2. Therefore, we should design a new framework of AdderNet for energy-efficient SISR models.
2 Learning Identity Mapping using AdderNet
Usually, an arbitrary super-resolution model using neural network learns the mapping from input LR image to HR image in an end-to-end manner. In addition to enhancing the high-frequency details, the overall texture and color information should be also maintained. Figure 2 illustrates the feature maps of different convolutional layers for a given LR image using VDSR . It can be seen that the difference between the input feature map and output feature map of each convolutional layer are very similar. This observation reveals that the identity mapping (\ie, ) in SISR task using deep learning method is very essential.
For the arbitrary low-resolution image , and an adder filter with its weight parameters . There exists no satisfying the following equation:
where denotes the adder operation defined in Eq. 2.
where .
Then, we can select proper values for each element in such that , we have:
Let , then we have:
Combining Eq. 5 and 6, we have , which is obviously impossible. In addition, the above proof can be easily extended to convolution layers in which the filter size of is often much smaller than that of the input data .
According to Theorem 1 and above analysis, the identity mapping cannot be directly learned using a one-layer adder neural network. To address this problem, we propose to refine the existing adder unit for adjusting the super-resolution task. In practice, we present a self-shortcut operation for each adder layer, \ie,
where is the weights of adder filters in the -th layer, and are the input data and output data, respectively. Since the output of Eq. 7 contains the input data itself, we can utilize it to approximate the identity mapping by reducing the magnitude of .
Different to the image recognition tasks, features in most of SISR problems should maintain a fixed size, \ie, the width and height of and are exactly the same. Thus, the new calculation as described in Eq. 7 can be embedded into most of conventional SISR models. In the following experiments, we will utilize Eq. 7 to replace most of convolutional layers whose output size is the same as that of the input size.
3 Learnable Power Activation
Besides the identity mapping, there is another important functionality of traditional convolution filters that cannot be easily ensured by adder filters. The goal of a SISR model is to enhance the details including color and texture information of the input low-resolution images. Therefore, the high-pass filter is also a very important component in most of existing SISR models. It can be found in Figure 2, the details of the input image are gradually enhanced when the network depth is increased.
Generally, natural images are composed of different-frequency information. For example, the background and a large area of grass are low-frequency information, where most of neighbor pixels are very closed. In contrast, edges of objects and some buildings are exactly high-frequency information for the given entire image. In practice, if we define an arbitrary image as the combination of high-frequency part and the low-frequency part as , an ideal high-pass filter used in the super-resolution task and other image processing problems could be defined as
which only preserves the high-frequency part of the input image. Wherein, is the convolution response of the high-frequency part. The above equation can help the SISR model removing redundant outputs and noise and enhancing the high-frequency details, which is also a very essential component in the SISR models.
Similarly, for the traditional convolution operation, the functionality of Eq. 8 can be directly implemented. For example, a high-pass filter can be used for removing any flat areas in . However, for the adder neural network as defined in Eq. 2, it is impossible to achieve the functionality as described in Eq. 8.
which means and leads to a contradiction.
According to the above theorem, the functionality of high-pass filter cannot be replaced by adder filter. Thus, the super-resolution process in the adder neural networks will involve more redundancy. To this end, we need to develop a new scheme for using adder neural networks to conduct the SISR tasks that can compensate this defect.
Admittedly, we can also add some parameters and filters to improve the capacity of the SR model using adder units, the save on energy and calculation will be reduced. Fortunately, Sharabati and Xi applied the Box-Cox transformation in the image denoising task and found that this transformation can achieve the similar functionality to that of the high-pass filters without adding massive parameters and calculations. Oliveira et al. further discussed the functionality of Box-Cox transformation in image super-resolution. In addition, a sign-preserving power law point transformation is also explored for emphasizing the areas with abundant details in the input image . Therefore, we propose a learnable power activation function to solve the defect of AdderNet and refine the output images, \ie,
where is the output features, is the sign function, is a learnable parameter for adjusting the information and distribution. When , the above activation function can enhance the contrast of output images and emphasize the high-frequency information. When , Eq. 11 can smooth all signals in the output image and remove artifacts and noise. In addition, the above function can be easily embedded into the conventional ReLU in any SISR models.
By exploiting these two methods described in Eq. 7 and Eq. 11, we can address the aforementioned problems on using adder networks for conducting the SISR task. Although there are some additional computations introduced, \eg, Eq. 7 needs a shortcut to maintain the information in the input data, and the learnable parameter in Eq. 11 leads to some additional multiplications. However, compared with the massive operations required by either convolutional layer or adder layer, they are very trivial. For example, the number of adder operations defined in Eq. 2 is about , and the additional computations required by Eq. 7 and Eq. 11 are both equal to . Considering that is usually a relatively large value in modern deep neural architectures (\eg, and ), the additional computation for each layer is over lower than that of the original method. In the next section, we will conduct extensive experiments to illustrate the superiority of the proposed method on both the visual quality and energy consumptions.
Experiments
In the above section, we have developed a series of new operations for establishing SISR models using adder neural networks. Here we will conduct experiments to verify the effectiveness of the proposed AdderSR networks.
To evaluate the performance of adder neural networks on super-resolution tasks, we select several benchmark image datasets to conduct the experiments. Following the setting of VDSR , 291 dataset is employed to train adder VDSR networks. It consists of 91 images from Yang et al. and 200 images from Berkeley Segmentation Dataset . In addition, DIV2K dataset is utilized to train adder EDSR networks in the following experiments. This dataset consists of 800 training images and 100 validation images. In order to compare with other state-of-the-art methods, four relatively small benchmarks are also selected including Set5, Set14, B100 and Urban100. The LR image is generated with bicubic downsampling. Super-resolution results are evaluated using both peak signal-to-noise ratio (PSNR ) and structure similarity index (SSIM ) on Y channel ( \ie, luminance) of YCbCr space.
Training setting.
We evaluate the performance of the proposed AdderSR networks using two famous neural architectures for super-resolution, \ie, VDSR and EDSR , which are shown extraordinary performance for generating images with high visual quality. Following the setting of the conventional AdderNet, we do not replace the first and the last convolutional layers in these networks, and the batch normalization is employed on each adder layer. To accelerate convergence speed of AdderSR networks, the learning rate for adder layers in our models is enlarged 10 times than that of convolutional layers in baselines. Specifically, the learning rates of adder layer and convolutional layer are initialized as and , respectively. The optimizer utilized in adder neural network for super-resolution task is ADAM due to its fast convergence speed. Hyper-parameters here are set as , and . Other settings such as patch size and data augmentation strategies are totally the same as those in baseline VDSR and EDSR.
Ablation study.
To ensure the performance of adder neural networks on the SISR tasks, we have thoroughly designed two new operations in Eq. 7 and Eq. 11. Here we will first conduct the detailed ablation study to illustrate their functionalities. In practice, the VDSR using convolutional layers is selected as baseline, and we replace all intermediate layers by adder layers as described in Eq. 2. Table 1 shows results of ablation experiments. Without Eq. 7 and Eq. 11, the PSNR of AdderSR is dB lower than that of conventional VDSR network on the average of four benchmark datasets. Such a PSNR value decline will increase the artifacts in the resulting high-resolution images. In Eq. 2, the self-shortcut makes it possible to optimize an identity mapping. Thus, the performance of AdderSR network using Eq. 2 can obtain an about 1.67 dB PSNR enhancement on average. In addition, Eq. 11 is proposed to emphasize the high-frequency information in intermediated features and reconstructed images. By embedding Eq. 11 into AdderSR network, the PSNR value can be further improved with about 0.24 dB on benchmark datasets. If both Eq. 7 and Eq. 11 are employed on AdderSR network, the performance can be improved with 1.91 dB on the average of four datasets. The final result of AdderSR network is pretty close to that of conventional VDSR. Detailed visualization results of these models can be found in the supplementary materials.
Feature Visulizations.
To have an explicit illustration on the functionalities of different components in our models using adder layers, we visualize the feature maps of Adder VDSR model 1, 3 and 4 reported in Table 1. From Figure 3, we can see that the feature maps of AdderSR network without Eq. 7 are quite different from the input image. The texture information such as human face, cars and trees in the input image are blured or distorted. In contrast, the features maps of AdderSR network using Eq. 7 preserve more details information than those of the model without Eq. 7. Moreover, it is obvious that the high-frequency information (\eg, edge, corner) is emphasized while a portion of low-frequency information is eliminated with the help of Eq. 11. That is a very important functionality for proceeding high-resolution images. In summary, by exploiting the proposed method, we can generate features with abundant texture and establish effective SISR models using only additions.
Overall Comparison.
After investigating the impact of different components of the proposed method, we then conduct the experiments using the benchmark VDSR and EDSR architectures with different scaling factors, \ie, , and to illustrate the superiority of the proposed AdderSR networks for the SISR problem. Table 2 shows the average PSNR and SSIM results of different models using both convolutional layers and adder units on Set5, Set14, B100, Urban100, respectively. It can be found in Table 2, our models using only additions in intermediate layers can also achieve comparable PSNR and SSIM values to those of their baselines with massive multiplications. The average PSNR gap between AdderSR models and baselines using convolutions over the four datasets is about 0.17 dB.
Besides the quantitative comparison, we also provide the comparison on the visual qualities of the same architecture using both additions and multiplications as shown in Figure 4. Actually, since the PSNR values of AdderSR networks and their baselines as shown in Table 2 are very closed, the output high-resolution images are of the similar visual quality. Particularly, the network using adder units can also effectively reconstruct the texture and color information. These results demonstrate that, besides the visual recognition task, we can also utilize additions to establish excellent neural networks for solving the SISR image processing problem. More visualization results of these models can be found in the supplementary materials.
Energy Consumptions.
As discussed in previous works , additions require much lower energy consumptions compared with multiplications. We further calculate the energy consumptions of different networks used in Table 2. In practice, most of values in recent SISR models are 32-bit floating numbers, and the energy consumptions for a 32-bit addition and multiplication are pJ and pJ, respectively. The amount of remaining multiplication operations in AdderSR network is extremely small compared with the FLOPs of the entire network. The detailed energy costs of the these networks are reported in Table 4, which is computed according to literature . Obviously, the proposed AdderSR method can reduce the energy cost for reconstructing a image by a fact of about 2.5. If we further quantize the weights and activations of these models shown in Table 2 to int8 values, we can obtain an about 3.8 reduction on the energy consumption using the proposed AdderSR with comparable performance. These results can be found in supplementary materials. We also conduct experiments using our methods on recent lightweight SR models (\ie, CARN ) and report the energy consumption v.s. performance comparison in Table 3, where CARN- , CARN- and CARN- are the models with , and channels after applying filter pruning, respectively. A-CARN denotes the model using AdderNets. We can achieve the comparable PSNR values on compact SISR architectures. The performance of A-CARN is about 0.4dB higher than that of CARN- with similar energy consumption (404GpJ) after pruning. This significant reduction on energy consumption will make these deep learning models portable on mobile devices.
Conclusions and Discussion
This paper investigates the single image super-resolution problem using AdderNets. Without changing the original neural architectures, we develop a new adder unit and a novel learnable power activation for addressing the defects in existing adder neural networks. The new AdderSR models can learn the functionalities of conventional identity mapping and high-pass filter, which are essential for providing images with high visual quality. Experimental results conducted on several benchmarks illustrate that, the proposed AdderSR networks can achieve the similar visual quality to that of their baselines using traditional convolution filters. Meanwhile, since additions are much cheaper than multiplications, we can reduce the energy consumptions of these models by about 2.5. Besides the image super-resolution task, the techniques in this paper can be well transferred to other image processing problems including denoising, deblurring, \etcFuture works will focus on more applications and low-bit quantization versions of these networks to achieve higher reduction on the energy consumption.