Dynamic Resolution Network

Mingjian Zhu, Kai Han, Enhua Wu, Qiulin Zhang, Ying Nie, Zhenzhong Lan, Yunhe Wang

Introduction

Deep convolutional neural networks (CNNs) have achieved remarkable success in various computer vision tasks, under the development of algorithms , computation power, and large-scale datasets . However, the outstanding performance is accompanied by large computational costs, which makes CNNs difficult to deploy on mobile devices. With the increasing demand for CNNs on real-world applications, it is imperative to reduce the computational cost and meanwhile maintain the performance of neural networks.

Recently, researchers have devoted much effort to model compression and acceleration methods, including network pruning, low-bit quantization, knowledge distillation, and efficient model design. Network pruning aims to prune the unimportant filters or blocks that are insensitive to model performance through a certain criterion . Low-bit quantization methods represent weights and activation in neural networks with low-bit values . Knowledge distillation transfers the knowledge of the teacher models to the student models to improve the performance . The efficient model design utilizes lightweight operations like depth-wise convolution to construct some novel architectures successfully. Orthogonal to those methods that usually focus on the network weights or architectures, Guo et.al. and Wang et.al. study the redundacy that exists in the input images. However, the resolutions of the input images in most of existing compressed networks are still fixed. Although deep networks are often trained using an uniform resolution (e.g., 224×\times 224 on the ImageNet), sizes and locations of objects in images are radically different. Figure 1 shows some samples that the required resolution for achieving the highest performance are different. For the given network architecture, the FLOPs (floating-number operations) of the network for processing image will be significantly reduced for images with lower resolution.

Admittedly, the input resolution is a very important factor that affects the computational costs and the performance of CNNs. For the same network, a higher resolution usually results in larger FLOPs and higher accuracy . In contrast, the model with a smaller input resolution has lower performance while the required FLOPs are also smaller. However, the shrink of input resolutions of deep networks provides us another potential to alleviate the computation burden of CNNs. To have an explict illustration, we first test some images under different resolutions with a pre-trained ResNet50 as shown in Figure 1 and count the minimum resolution required to give the correct prediction for each sample. In practice, “easy” samples, such as the panda with obvious foreground, can be classified correctly in both low and high resolution, and “hard” samples, such as the damselfly whose foreground and background are tangled can only be predicted accurately in high resolution. This observation indicates that a larger proportion of images in our datasets can be efficiently processed by reducing their resolutions. On the other hand, it is also compatible with the human perception system , i.e., some samples can be understood easily just in blurry mode while the others need to be seen in clear mode.

In this paper, we propose a novel dynamic-resolution network (DRNet) which dynamically adjusts the input resolution of each sample for efficient inference. To accurately find the required minimum resolution of each image, we introduce a resolution predictor which is embedded in front of the entire network. In practice, we set several different resolutions as candidates and feed the image into the resolution predictor to produce a probability distribution over candidate resolutions as the output. The network architecture of the resolution predictor is carefully designed with negligible computational complexity and trained jointly with classifier for recognition in an end-to-end fashion. By exploiting the proposed dynamic resolution network inference approach, we can excavate the reduncancy of each image from its input resolution. Thus, computational costs of easy samples with lower resolutions can be saved, and the accuracy for hard samples can also be preserved by maintaining higher resolutions. Extensive experiments on the large-scale visual benchmarks and the conventional ResNet architectures demonstrate the effectiveness of our proposed method for reducing the overall computational costs with comparable network performance.

Related Works

Although deep CNN models have shown excellent accuracy, they often contain millions of parameters and FLOPs. Thus the model compression techniques are becoming a research hostpot for reducing the computational costs of CNNs. Here, we revisit existing works in two parts, i.e., static model compression and dynamic model compression.

We classify model compression methods that are not instance-aware as the static. Group-wise Convolution (GWC), Depth-wise Convolution and Point-wise Convolution are widely used for efficient model design, such as MobileNet , ResNeXt , and ShuffleNet . Revealing pattern redundancy among feature maps, Han proposes to generate more feature maps from intrinsic ones through some cheap operations and Zhang et.al. adopts relatively heavy computation to extract intrinsic information while tiny hidden details are processed with some light-weight operations. The methods above achieve model acceleration to some extent. However, they treat all input samples equally, whereas the difficulty for CNN models or humans to recognize each sample is unequal. So instance-aware model compression can be further explored.

Dynamic model compression takes the unequal difficulty of each sample into consideration. Huang proposes multi-scale dense networks with multiple classifiers to allocate uneven computation across "easier" and "harder" inputs. Wu et.al. introduces BlockDrop which learns to dynamically execute the necessary layers so as to best reduce total computation. Veit et.al. proposes ConvNet-AIG to adaptively define their network topology conditioned on the input image. Except for dynamic adjustment on model architectures, recent works pay more attention to input images. Verelst et.al. proposes a small gating network to predict pixel-wise masks determining the locations where dynamic convolutions are evaluated. Uzkent et.al. proposes PatchDrop to dynamically identify when and where to use high-resolution data conditioned on the paired low-resolution images. Yang et.al. dynamically utilize sub-networks of the base network to process images with different resolutions. Wang et.al. proposes GFNet with patch proposal networks that strategically crop the minimal image regions to obtain reliable predictions. There exists many works on multiple resolutions and dynamic mechanisms. ELASTIC uses different scaling policies for different instances and it learns from the data how to select the best policy. Hydranets chooses different branches for different inputs by a gate and aggregates their outputs with a combiner. RS-Net utilizes private BNs, shared convolutions, and fully-connected layer and to train input images with different resolutions.

Different from dynamic adjustment on model architectures and dynamic modification on input images with reinforcement learning, we consider the whole image and propose a resolution predictor to dynamically choose the performance-sufficient and cost-efficient resolution for a single model to obtain reliable prediction with end-to-end training.

Approach

In this section, we first introduce the overall framework of the proposed Dynamic Resolution Network (DRNet), then describe the resolution predictor, resolution-aware BN, and optimization algorithm in detail, respectively.

Inspired by the fact that different sample requires different resolution to achieve the least accurate prediction, we propose to develop an instance-aware resolution selection approach for a single large classifier network. As shown in Figure 2, the proposed method mainly consists of two components. The first is the large classifier network with both high performance and expensive computational costs, such as the classical ResNet and EfficientNet . The other is a resolution predictor for finding the minimal resolution so that we can adjust the input resolution of each image to have a better trade-off on the accuracy and efficiency. For an arbitrary input image, we first forecast its suitable resolution rr using the resolution predictor. Then, the large classifier will take the resized image as inputs and the required FLOPs will be reduced significantly when rr is lower than that of the origional resolution. To achieve better performance, the resolution predictor and the base network are optimized end-to-end during training.

The resolution predictor is designed as a CNN-based preprocessing operation before input samples are fed to the large base network. It’s from the inspiration that our well-trained large models can also predict a relative amount of samples correctly though they are in relatively small resolutions while large amounts of computation cost can be saved. On the one hand, the goal of the resolution predictor is to find an appropriate instance-aware resolution by inferring a probability distribution over candidate resolutions. Note that there are a vast number of candidates from 1×\times1 to 224×\times224, which makes it difficult, also meaningless, for the resolution predictor to explore such a long-range of resolutions. As a simplification strategy and practical requirement, we choose mm resolution candidates r1,r2,⋯ ,rmr_{1},r_{2},\cdots,r_{m} to shrink the exploration range. On the other hand, we have to keep the model size of the proposed resolution predictor as small as possible since it will bring extra FLOPs, otherwise, it becomes impractical to implement such a module if its extra introduced computation exceeds the saved one from the low resolution. In this spirit, we design the resolution predictor with a few convolutional layers and fully-connected layers to complete a resolution classification task. Then the preprocessing of the proposed resolution predictor R(⋅)R(\cdot) can be given as follows:

where the Gumbel-Softmax trick will be described in the next subsection. For validation, the resolution predictor makes decisions first then the input with the selected resolution only is fed to the large classifier in the normal way as shown in the left part of Figure 2.

Resolution-aware BN.

Our framework is proposed to use only a single large classifier for the sake of storage pressure and loading latency, which results in that the single classifier has to process multi-resolution inputs and raises two problems. One obvious problem is that the first fully-connected layer will fail to work with a different input resolution and can be solved with global average pooling. Thus we can process multiple resolutions in one single network. The other hidden problem exists in Batch Normalization (BN) layers. BN is used to make deep models converge faster and more stable through channel-wise normalization of the input layer by re-centering and re-scaling. However, activation statistics including means and variances under different resolutions are incompatible . Using shared BNs under multiple resolutions leads to lower accuracy in our experiments as shown in section 4.3. Since the batch normalization layer contains a negligible amount of parameters, we propose resolution-aware BNs as shown in Figure 2. We decouple the BN for each resolution and choose the corresponding BN layer to normalize the features:

where ϵ\epsilon is a small number for numerical stability, μi{\mu}_{i} and σi{\sigma}_{i} are private averaged mean and variance from the activation statistics under separate resolutions; βi{\beta_{i}} and γi{\gamma_{i}} are private learnable scale weights and bias. Since shared convolutions are insensitive to performance, the overall adjustment for the original large classifier is shown in the right part of Figure 2.

2 Optimization

The proposed framework is optimized to perform instance-aware resolution selection for inputs of a single large classifier with end-to-end training. The loss function and Gumbel softmax trick are described in the following.

The base classifier and the resolution predictor are optimized jointly. The loss function includes two parts: the cross-entropy loss for image classification and a FLOPs constraint regularization to restrict the computation budget.

The Cross-Entropy loss H\mathcal{H} is performed between y^\hat{y} and target label yy as follows:

The gradients from the loss LceL_{ce} are back-propagated to both the base classifier and the resolution predictor for optimization.

If we use the Cross-Entropy loss only, the resolution predictor will converge to a sub-optimal point and tend to select the largest resolution because samples with the largest resolution correspond to relatively lower classification loss generally. Although the classification confidence of the low-resolution image is relatively lower, the prediction can be correct and requires fewer FLOPs. In order to reduce the computational cost and balance the different resolution selection, we propose a FLOPs constraint regularization to guide the learning of resolution predictor:

Finally, the overall loss is the weighted summation over the classification loss and the FLOPs constraint regularization term:

where η\eta is a hyper-parameter to match the magnitude of LceL_{ce} and LregL_{reg}.

Gumbel Softmax Trick.

Since there exists an non-differentiable problem in the process from the resolution predictor’s continues outputs to discrete resolution selection, we adopt Gumbel Softmax trick to make discrete decision differentiable during the back-propagation. In Eq. 1, the resolution predictor gives the probabilities for the resolution candidates pr=[pr1,pr2,...,prm]p_{r}=[p_{r_{1}},p_{r_{2}},...,p_{r_{m}}]. Then the discrete candidate resolution selections can be drawn using:

where gjg_{j} is Gumbel noise obtained through two log⁡\log operation applied on i.i.d samples uu drawn from a uniform distribution as follows:

During training, the derivative of the one-hot operation is approximated by Gumbel softmax function which is both continuous and differentiable:

where τ\tau is the temperature parameter. The introduction of Gumbel noise has two positive effects. On the one hand, it will not influence the highest entry of the original categorical probability distribution. On the other hand, it makes the gradient approximation from discrete hardmax to continuous softmax more fluent. By this straight-through Gumbel softmax trick, we can optimize the overall framework end-to-end.

Experiments

To show the effectiveness of our proposed method, in this section, we conduct experiments on the small-scale ImageNet-100 and large-scale ImageNet-1K with classic large classifier networks, including ResNet and MobileNetV2 , where we replace their single batch normalization layer(BN) with resolution-aware BNs and add the proposed resolution predictor to guide the resolution selection.

ImageNet-1K dataset (ImageNet ILSVRC2012) is a widely-used benchmark to evaluate the classification performance of neural networks, which consists of 1.28M training images and 50K validation images in 1K categories. ImageNet-100 is a subset of the ImageNet ILSVRC2012, whose training set is random selected from the original training set and consists of 500 instances of 100 categories. The validation set is the corresponding 100 categories of the original validation set. The categories of ImageNet-100 is provided in supplementary materials. For the license of ImageNet dataset, please refer to http://www.image-net.org/download.

Experimental Settings.

For data augmentation during training for both ImageNet-100 and ImageNet-1K, we follow the scheme as in including randomly cropping a patch from the input image and resizing to candidate resolutions with the bilinear interpolation followed by random horizontal flipping with probability 0.5. For data processing during validation, we first resize the input image into 256×256256\times 256 and then crop the center 224×224224\times 224 part. The details of the resolution predictor are provided in supplementary materials. For both datasets, we firstly employ the images of different resolutions to pre-train a model without the resolution predictor. The losses of each resolution are summed up for optimization. Then we add a designed predictor to the model and conduct finetuning. Optimization is performed using SGD (mini-batch stochastic gradient descent) and learning rate warmup is applied for the first 3 epochs. In the pretraining stage, the model is trained with total epochs 70, batch-size 256, weight decay 0.0001, momentum 0.9, initial learning rate 0.1 which decays a factor of 10 every 20 epochs. We adopt a similar training scheme in finetuning stage. The total epochs are 100 with the learning rate decaying a factor of 10 every 30 epochs. We adopt 1×\times learning rate to finetune the large classifier and 0.1×\times learning rate to train the resolution predictor from scratch. The framework is implemented in Pytorch on NVIDIA Tesla V100 GPUs.

2 ImageNet-100 Experiments

We conduct small-scale experiments on ImageNet-100 to guide the resolution selection for large classifiers ResNet-50. The resolution predictor is designed as a 4-stage residual network with input resolution 128×128128\times 128 where each stage contains one residual basic block, which consumes about 300 million FLOPs. For candidate resolutions of the large classifier, we choose resolutions of [224×224224\times 224, 168×168168\times 168, 112×112112\times 112] and we denote them as for simplicity. Thus the last fully-connected layers of the resolution predictor contain three neurons. We only replace each batch normalization layer in ResNet-50 with three optional resolution-aware batch normalization layers, change the last average pooling layer to an adaptive average pooling layer, and then integrate the resolution predictor to form the overall framework. Experiment results are shown in Table 2.

For the calculation of the average FLOPs in Table 2, we sum up the FLOPs of each sample under the predicted resolution, take the extra FLOPs introduced by the resolution predictor into account and finally take the average over the whole validation set. From the results in Table 2, we can see our dynamic-resolution ResNet-50 obtains about 17% reduction of average FLOPs while gains 4.0% accuracy increase with the candidate resolutions . When we tune the hyperparameters (i.e., η\eta and α\alpha) in FLOPs constraint regularization, the dynamic-resolution ResNet-50 obtains about 32% FLOPs reduction and achieve 1.8% increase in accuracy. We also extend the range of resolutions (i.e., ) for fully exploration, especially the lower resolution. We can see that our DRNet still performs better than the baseline model. Setting a larger α\alpha can even obtain 44% FLOPs reduction with performance increase, which is shown in Table 2.

3 Ablation Study

To form the dynamic resolution network, we propose two adjustments, 1) replacing each BN layer with resolution-aware BNs; 2) proposing a FLOPs balance regularizer. Here we conduct ablation studies on ImageNet-100 to investigate the influence of each part.

Here we compare the results where the large classifier ResNet-50 is equipped with resolution-aware BN or not. From Table 2, we can see that resolution-aware BNs obtain extra one more points for ResNet-50 with similar computational cost, which demonstrates that we need to normalize feature maps with different resolutions separately, thus their activation statistics can be more accurate.

Influence of FLOPs Constraint Regularization.

Here we explore the influence of the penalty factor η\eta and target FLOPs value α\alpha in the FLOPs constraint regularization as shown in Table 2. We first fix η\eta as 0.2 and tune α\alpha from 2.0 to 3.5. We can see that the average FLOPs increase from 2.3G to 3.5G gradually, and the accuracy also increase consequently. As for η\eta, we fix α=3.0\alpha=3.0 and tune η\eta in the range of [0.1, 1]. A larger penalty factor leads to lower FLOPs and accuracy. That is to say when selecting images dynamically, the resolution predictor with lower penalty η\eta tends to choose the larger resolution, where the effect of the balance regularizer is relatively weaker.

4 ImageNet-1K Experiments

We conduct large-scale experiments with DR-ResNet-50 on ImageNet-1K as shown in Table 3. When we set the candidate resolutions as , our DR-ResNet-50 also outperforms the baseline by 1.4 percentage with 10% FLOPs reduced. Similar to the results in Table 2, the effectiveness of FLOPs constraint regularization is also verified in Table 3, where the FLOPs drop with larger α\alpha. Our DRNet focuses on input resolution and keeps the structure of the large classifier almost unchanged, thus the parameters of our DRNet-equipped models are more than the original large classifier due to the introduction of the resolution predictor. In other words, since our DRNet is orthogonal to those architecture compression methods, careful combinations of the two methods would make a more compact result.

We also compare DR-ResNet-50 with other representative model compression methods to verify the superiority of the proposed method. The compared methods include Sparse Structure Selection (SSS) , Versatile Filters , PFP , and C-SGD . As shown in Table 4, our DR-ResNet-50 achieves better performance than other methods with similar FLOPs.

Effect of Dynamic Resolution.

To evaluate the effect of the proposed dynamic resolution mechanism, we compare DR-ResNet-50 with randomly selected resolution. We repeat the random selection 3 times and report their accuracies and FLOPs in Imagenet-1K dataset. From the results in Table 3, DRNet shows much better performance than random baseline, indicating the effectiveness of dynamic resolution.

On-device Acceleration of Dynamic Resolution.

In Figure LABEL:fig:latency, we demonstrate the practical accelerations of our DR-ResNet-50, which are obtained by measuring the forward time on a Intel(R) Xeon(R) Gold 6151 CPU. We directly set the batch size as 1 and the input resolution for the resolution predictor is 128. The candidate resolutions are 224, 168, and 112. We average the test time in ImageNet-1K val set. Our model substantially outperforms ResNet-50 by a significant margin.

MobileNetV2 Results.

We also test our method on a representative lightweight neural network, i.e.MobileNetV2 . We set the candidate resolution as . The training setting of MobileNetV2 follows that in the original paper for a fair comparison. To reduce the FLOPs of the resolution predictor, we replace the residual block with an inverted residual block and set the input size as 64×\times64. From the results in Table 6, we can see that DRNet achieves 72.7% top-1 accuracy with fewer computational costs.

5 Visualization

The prediction results of the resolution predictor are visualized in Figure 4. The first four samples with obvious foreground which occupy most of the whole image are predicted to use 112×112112\times 112 resolution in high confidence. The middle three whose foreground is a little blurred are predicted to select 168×168168\times 168 resolution. The last three samples’ hidden foregrounds nearly blend with the background, thus the largest resolution are selected. Although the “easy” and “hard” examples may be different for humans and machines, these results are compatible with the human perception system.

Conclusion

In this paper, we reveal that different sample acquires different resolution threshold to achieve the least accurate prediction. Thus large amounts of computation cost can be saved for some easier samples under lower resolutions. To make CNNs predict efficiently, we propose a novel dynamic resolution network to dynamically choose the performance-sufficient and cost-efficient resolution for each input sample. Then the input is resized to the predicted resolution and fed to the original large classifier, in which we replace each BN layer with resolution-aware BNs to accommodate the multi-resolution input. The proposed method is decoupled with the network architecture and can be generalized to any network. Extensive experiments on various networks demonstrate the effectiveness of DRNet.

Appendix A Appendix

We utilize the basic block of resnet to construct the predictor, i.e., 4 basic blocks are stacked to form the predictor network in our paper. Here we build 4 predictor architectures with fewer parameters and FLOPs and we compare their performances as follows. We can see that the predictor with fewer flops leads to slight accuracy degradation.

(1) Predictor-Architecture-1: The original predictor in our paper. The parameters of the first convolution are Conv2d (3, 64, kernel_size=(7, 7), stride=(2, 2), padding=(3, 3), bias=False).

(2) Predictor-Architecture-2: We reduce the blocks of the predictor in (1) to construct a new predictor. We retain only 2 blocks.

(3) Predictor-Architecture-3: We increase the stride of the first convolution in (2) to 4. Thus, the parameters of the first convolution are Conv2d (3, 64, kernel_size=(7, 7), stride=(4, 4), padding=(3, 3), bias=False).

(4) Predictor-Architecture-4: We construct the predictor with only two convolutions.

The details of these predictors are shown as follows:

A.2 ImageNet-100 Categories.

References