DO-Conv: Depthwise Over-parameterized Convolutional Layer
Jinming Cao, Yangyan Li, Mingchao Sun, Ying Chen, Dani Lischinski, Daniel Cohen-Or, Baoquan Chen, Changhe Tu
Introduction
Convolutional Neural Networks (CNNs) are capable of expressing highly complicated functions, and have shown great success in addressing many classical computer vision problems, such as image classification, detection and segmentation. It has been widely accepted that increasing the depth of a network by adding linear and non-linear layers together can increase the network’s expressiveness and boost its performance. On the other hand, adding extra linear layers only is not as commonly considered, especially when the additional linear layers result in an over-parameterizationThe meaning of “over-parameterization” is overloaded in the community. We opt to use this term and its meaning following . The term is also used in the community for describing the case where the number of parameters in a model exceeds the size of the training dataset, such as that in . — a case where the composition of consecutive linear layers may be represented by a single linear layer with fewer learnable parameters.
Though over-parameterization does not improve the expressiveness of a network, it has been proven as means of accelerating the training of deep linear networks, and shown empirically to speedup the training of deep non-linear networks . These findings suggest that, while much work has been devoted to the quest for novel network architectures, over-parameterization has considerable unexplored potential in benefiting existing architectures.
In this work, we propose to over-parameterize a convolutional layer by augmenting it with an “extra” or “over-parameterizing” component: a depthwise convolution operation, which convolves separately each of the input channels. We refer to this depthwise over-parameterized convolutional layer as DO-Conv, and show that it can not only accelerate the training of various CNNs, but also consistently boost the performance of the converged models.
One notable advantage of over-parameterization, in general, is that the multi-layer composite linear operations used by the over-parameterization can be folded into a compact single layer representation after the training phase. Then, only a single layer is used at inference time, reducing the computation to be exactly equivalent to a conventional layer.
We show with extensive experiments that using DO-Conv boosts the performance of CNNs on many tasks, such as image classification, detection, and segmentation, merely by replacing the conventional convolution with DO-Conv. Since the performance gains are introduced with no increase in inference computations, we advocate DO-Conv as an alternative to the conventional convolutional layer.
Related Work
The performance of CNNs is highly correlated with their architectures, and a line of notable architectures have been proposed over the recent years, such as AlexNet , VGG , GoogLeNet , ResNet , and MobileNets . Our work is orthogonal to the quest for novel architectures, and can be combined with existing ones to boost their performance.
Convolutional layers are the core building blocks of CNNs. Thus, an improvement of these core building blocks can often lead to boosting the performance of CNNs. Several alternatives to the classical convolutional layer have been proposed, which offer improved feature learning capability and/or efficiency . Our work may be viewed as a contribution along this line.
Arora et al. studied the role of over-parameterization in gradient descent based optimization. They have proven that over-parameterization of fully connected layers can accelerate the training of deep linear networks, and shown empirically that it can also accelerate the training of deep non-linear networks. There are multiple ways to over-parameterize a layer. The kernel of a convolutional layer has both channel and spatial axes, thus the over-parameterization of a convolutional layer can be more versatile than that of a fully connected layer. ExpandNets proposes to over-parameterize the convolution kernel over the input and output channel axes. The advantage of over-parameterization is demonstrated in only on the training of compact CNNs, whose performance is no match to that of the mainstream CNNs. Our method over-parameterizes the convolution kernel over its spatial axes, and the advantage is demonstrated on several commonly used CNN architectures. In ACNet the convolutional layer is replaced with an Asymmetric Convolution Block. This approach is, in essence, an over-parameterization of the convolution kernel over the center row and column of its spatial axes. In fact, ACNet may be viewed as a special case of our over-parameterization approach, which over-parameterizes the entire kernel.
Over-parameterized layers introduce “extra” linear transformations without increasing the expressiveness of the network, and once their parameters have been learned, these “extra” transformations are folded in the inference phase. In this sense, normalization layers, such as batch normalization and weight normalization , are quite similar to over-parameterized layers. Normalization layers have been widely used for improving the training of CNNs, while the reason for their success is still being actively studied . It is intriguing to study whether or not the effectiveness of over-parameterization and normalization in CNNs may be attributed to the same underlying reasons.
Method
Conventional convolutional layer.
Depthwise convolutional layer.
Depthwise over-parameterized convolutional layer (DO-Conv)
DO-Conv is an over-parameterization of convolutional layer.
Training and inference of CNNs with DO-Conv.
Training efficiency and composition choice of DO-Conv.
The two ways for computing DO-Conv are mathematically equivalent, but have different training efficiency. The number of multiply and accumulate operations (MACC), is often used for measuring the amount of computation and serves as an indicator of the efficiency. The MACC cost of feature and kernel composition, when they are applied on an feature map in (assuming ), can be calculated as follows:
DO-Conv and depthwise separable convolutional layer [7].
The feature composition of DO-Conv (Figure 3-a) is equivalent to applying a depthwise convolutional layer on an input feature map, yielding an intermediate feature map in channels, and then applying a convolutional layer on the intermediate feature map. This process is exactly same as that of a depthwise separable convolutional layer, where the convolution is often referred to as pointwise convolution. However, the motivation of DO-Conv and depthwise separable convolution is quite different. Depthwise separable convolution is introduced as an approximation and alternative to a conventional convolution for saving MACC to ease the deployment of CNNs especially on edge devices, thus it requiresIn earlier versions of Tensorflow , setting in the depthwise separable convolutional layer is considered as an error, though this constraint has been removed since commit 1822073. , and is probably the choice most widely used in practice, for example in Xception and MobileNets .
Training depthwise separable convolutional layer with kernel composition.
The equivalence between feature and kernel composition holds not only for (DO-Conv), but also for (depthwise separable convolution). We have shown that for DO-Conv, kernel composition often saves MACC and memory, compared with feature composition. For depthwise separable convolution, it is easy to see that feature composition is more economical. While the MACC is highly correlated with the running speed on edge devices that are often computation capability bounded, this might not be the case on NVIDIA GPUs, as they are more memory access bounded. Depthwise separable convolution has a significantly lower MACC cost than a conventional convolution, but it often runs slower on NVIDIA GPUs, which the CNNs are often trained on. To speedup the training of depthwise separable convolution layers on NVIDIA GPUs, kernel composition can be used, after which feature composition can be used for the deployment on edge devices. This handy trick is not the focus of this paper, but a by-product of the feature and kernel composition equivalence that may greatly interest many deep learning practitioners.
Depthwise over-parameterized depthwise/group convolutional layer (DO-DConv/DO-GConv).
Depthwise over-parameterization can not only be applied over a conventional convolution to yield DO-Conv, but also be applied over a depthwise convolution, which leads to DO-DConv. Following the same principle we used for establishing DO-Conv, as shown in Figure 4, DO-DConv also can be computed in two mathematically equivalent ways as:
Experiments
Comparison protocol.
The performance of a CNN can be affected by a wide range of factors. To demonstrate the effectiveness of DO-Conv, we take several notable CNNs as baselines, and merely replace the non-pointwise convolutional layers with DO-Conv, without any change to other settings. In other words, the replacement is the one and only difference between the baseline and our method. This guarantees that the observed performance change is due to the application of DO-Conv, but not other factors. Furthermore, this also means that no hyper-parameter is tuned to favor DO-Conv. Following this protocol, we demonstrate the effectiveness of DO-Conv on image classification, semantic segmentation and object detection tasks on several benchmark datasets.
1 Image Classification
We conducted image classification experiments on CIFAR and ImageNet , with a set of notable architectures including ResNet-v1 (including ResNet-v1b, which modifies ResNet-v1 by setting the stride at the layer of a bottleneck block), Plain (same as ResNet-v1, but without skip links), ResNet-v2 , ResNeXt , MobileNet-v1 , -v2 and -v3 .
The experiments on CIFAR follow the same settings as those in and the results are shown in Table 1, where the “DO-Conv” rows show the performance change relative to the baselines. All the results on CIFAR dataset reported in this paper are the averaged accuracy of the last five epochs over five runs. We can observe that DO-Conv brings a promising improvement over the baselines on relatively shallower networks.
The experiments on ImageNet follow the same settings as those in the model zoo of GluonCV . We implemented DO-Conv into GluonCV (commit bbe4166), such that it is convenient to make sure DO-Conv is the only change over baselines, since all of the other settings are not touched. We base the experiments on GluonCV since it provides not only a wide variety of high performance CNNs, but also the their training procedures. We consider GluonCV highly reproducible, but still, to exclude clutter factors as much as possible, we train these CNNs as baseline ourselves, and compare DO-Conv versions with them, while reporting the performance provided by GluonCV as reference. The results of DO-Conv and DO-DConv/DO-GConv are summarized in Table 2 and Table 3, respectively. It is clear that DO-Conv consistently boosts the performance of various baselines on ImageNet classification.
2 Semantic Segmentation
We also conducted semantic segmentation experiments on the PASCAL VOC and the Cityscapes datasets with GluonCV, following the same comparison protocol as that used for the ImageNet classification task. Differently from image classification, the training of a CNN model for image segmentation often has two stages: the “Backbone” and “Segmentation”. The first stage pre-trains a backbone model on the ImageNet classification task, and the second stage fine-tunes the backbone for the segmentation task. DO-Conv can be used in one or both of the stages. We use Deeplabv3 with ResNet-50 and ResNet-100 as the backbones for Cityscapes and PASCAL VOC datasets, respectively, and summarize the results in Table 4. Note that the delta in both the second and third rows are relative to the first row. We can observe that DO-Conv consistently boosts the performance, either when used only in the second stage, or in both stages.
3 Object Detection
We evaluated DO-Conv for object detection task on the COCO dataset using Faster R-CNN with ResNet-50 backbone, again with GluonCV, following the same aforementioned comparison protocol. The results are summarized in Table 5. Similarly to segmentation, the detection task has two stages: the “Backbone” and “Detection”. Note that the delta in both the second and third rows are relative to the first row. We can observe that using DO-Conv in the “Detection” stage only does not improve the overall performance. Yet, when it is used in both stages, an obvious improvement is achieved. Remember that no hyper-parameter tuning was done when making these comparisons, and all the experiments stick to the hyper-parameters tuned for the baselines.
4 Visualizations
We show the train and validation curves of baseline and DO-Conv over 120 epochs with ResNet-v1b in different depths in Figure 5. The hyper-parameters of ResNet in GluonCV are different from those used in the original ResNet work . One notable difference is that the learning rate is gradually decayed to zero at the end of training, and the training converges to a significantly better performance than that reported in . It is clear that the training of DO-Conv not only converges faster, but also converges to lower errors. While faster convergence due to over-parameterization has been reported before , to the best of our knowledge, we are the first to report convergence to lower errors on mainstream architectures.
5 Ablation Studies
We show the results of replacing the convolutional layers in subsets of ResNet stages with DO-Conv in Table 6. For ResNet-v2-20 on CIFAR-100, the accuracy consistently improves when more DO-Conv are used. For ResNet-v1b-50 on ImageNet, the effect is more complicated, as the use of DO-Conv in certain stages may result in a performance drop. The effectiveness of DO-Conv in the first stage is consistent in both cases.
The kernels in CNNs are usually randomly initialized when training the networks from scratch. Since the role of the over-parameterization kernel is different than that of other kernels, we opt to initialize it as identity. However, we found that even if it is randomly initialized (with He initialization ), it can still boost the performance, though the gain is smaller, compared to identity initialization, as shown in Table 7.
Conclusions and Future work
DO-Conv, a depthwise over-parameterized convolutional layer, is a novel, simple and generic way for boosting the performance of CNNs. Beyond the practical implications of improving training and final accuracy for existing CNNs, without introducing extra computation at the inference phase, we envision that the unveiling of its advantages could also encourage further exploration of over-parameterization as a novel dimension in network architecture design.
In the future, it would be intriguing to get a theoretical understanding of this rather simple means in achieving the surprisingly non-trivial performance improvements on a board range of applications. Furthermore, we would like to expand the scope of applications where these over-parameterized convolution layers may be effective, and learn what hyper-parameters can benefit more from it.
Broader Impact
We believe that the impact of our work is mainly to improve the performance of existing CNN models on a variety of computer vision tasks. It could also assist in developing CNN-based solutions to tasks which have not yet been tackled in this manner. We do not believe that our work has any ethical or social implications beyond those stated above.