BAM: Bottleneck Attention Module
Jongchan Park, Sanghyun Woo, Joon-Young Lee, In So Kweon
Introduction
Deep learning has been a powerful tool for a series of pattern recognition applications including classification, detection, segmentation and control problems. Due to its data-driven nature and availability of large scale parallel computing, deep neural networks achieve state-of-the-art results in most areas. Researchers have done many efforts to boost the performance in various ways such as designing optimizers [Zeiler(2012), Kingma and Ba(2014)], proposing adversarial training scheme [Goodfellow et al.(2014)Goodfellow, Pouget-Abadie, Mirza, Xu, Warde-Farley, Ozair, Courville, and Bengio], or task-specific meta architecture like 2-stage architectures [Ren et al.(2015)Ren, He, Girshick, and Sun] for detection.
A fundamental approach to boost performance is to design a good backbone architecture. Since the very first large-scale deep neural network AlexNet [Krizhevsky et al.(2012)Krizhevsky, Sutskever, and Hinton], various backbone architectures such as VGGNet [Simonyan and Zisserman(2014)], GoogLeNet [Szegedy et al.(2015)Szegedy, Liu, Jia, Sermanet, Reed, Anguelov, Erhan, Vanhoucke, and Rabinovich], ResNet [He et al.(2016b)He, Zhang, Ren, and Sun], DenseNet [Huang et al.(2016a)Huang, Liu, Weinberger, and van der Maaten], have been proposed. All backbone architectures have their own design choices, and show significant performance boosts over the precedent architectures.
The most intuitive way to boost the network performance is to stack more layers. Deep neural networks then are able to approximate high-dimensional function using their deep layers. The philosophy of VGGNet [Simonyan and Zisserman(2014)] and ResNet [He et al.(2016a)He, Zhang, Ren, and Sun] precisely follows this. Compared to AlexNet [Krizhevsky et al.(2012)Krizhevsky, Sutskever, and Hinton], VGGNet has twice more layers. Furthermore, ResNet has 22x more layers than VGGNet with improved gradient flow by adopting residual connections. GoogLeNet [Szegedy et al.(2015)Szegedy, Liu, Jia, Sermanet, Reed, Anguelov, Erhan, Vanhoucke, and Rabinovich], which is also very deep, uses concatenation of features with various filter sizes at each convolutional block. The use of diverse features at the same layer shows increased performance, resulting in powerful representation. DenseNet [Huang et al.(2016a)Huang, Liu, Weinberger, and van der Maaten] also uses the concatenation of diverse feature maps, but the features are from different layers. In other words, outputs of convolutional layers are iteratively concatenated upon the input feature maps. WideResNet [Zagoruyko and Komodakis(2016)] shows that using more channels, wider convolutions, can achieve higher performance than naively deepening the networks. Similarly, PyramidNet [Han et al.(2017)Han, Kim, and Kim] shows that increasing channels in deeper layers can effectively boost the performance. Recent approaches with grouped convolutions, such as ResNeXt [Xie et al.(2016)Xie, Girshick, Dollár, Tu, and He] or Xception [Chollet(2016)], show state-of-the-art performances as backbone architectures. The success of ResNeXt and Xception comes from the convolutions with higher cardinality which can achieve high performance effectively. Besides, a practical line of research is to find mobile-oriented, computationally effective architectures. MobileNet [Howard et al.(2017)Howard, Zhu, Chen, Kalenichenko, Wang, Weyand, Andreetto, and Adam], sharing a similar philosophy with ResNeXt and Xception, use depthwise convolutions with high cardinalities.
Apart from the previous approaches, we investigate the effect of attention in DNNs, and propose a simple, light-weight module for general DNNs. That is, the proposed module is designed for easy integration with existing CNN architectures. Attention mechanism in deep neural networks has been investigated in many previous works [Mnih et al.(2014)Mnih, Heess, Graves, et al., Ba et al.(2014)Ba, Mnih, and Kavukcuoglu, Bahdanau et al.(2014)Bahdanau, Cho, and Bengio, Xu et al.(2015)Xu, Ba, Kiros, Cho, Courville, Salakhudinov, Zemel, and Bengio, Gregor et al.(2015)Gregor, Danihelka, Graves, Rezende, and Wierstra, Jaderberg et al.(2015a)Jaderberg, Simonyan, Zisserman, et al.]. While most of the previous works use attention with task-specific purposes, we explicitly investigate the use of attention as a way to improve network’s representational power in an extremely efficient way. As a result, we propose “Bottleneck Attention Module” (BAM), a simple and efficient attention module that can be used in any CNNs. Given a 3D feature map, BAM produces a 3D attention map to emphasize important elements. In BAM, we decompose the process of inferring a 3D attention map in two streams (Fig. 2), so that the computational and parametric overhead are significantly reduced. As the channels of feature maps can be regarded as feature detectors, the two branches (spatial and channel) explicitly learn ‘what’ and ‘where’ to focus on.
We test the efficacy of BAM with various baseline architectures on various tasks. On the CIFAR-100 and ImageNet classification tasks, we observe performance improvements over baseline networks by placing BAM. Interestingly, we have observed that multiple BAMs located at different bottlenecks build a hierarchical attention as shown in Fig. 1. Finally, we validate the performance improvement of object detection on the VOC 2007 and MS COCO datasets, demonstrating a wide applicability of BAM. Since we have carefully designed our module to be light-weight, parameter and computational overheads are negligible.
Contribution. Our main contribution is three-fold.
We propose a simple and effective attention module, BAM, which can be integrated with any CNNs without bells and whistles.
We validate the design of BAM through extensive ablation studies.
We verify the effectiveness of BAM throughout extensive experiments with various baseline architectures on multiple benchmarks (CIFAR-100, ImageNet-1K, VOC 2007 and MS COCO).
Related Work
A number of studies [Itti et al.(1998)Itti, Koch, and Niebur, Rensink(2000), Corbetta and Shulman(2002)] have shown that attention plays an important role in human perception. For example, the resolution at the foveal center of human eyes is higher than surrounding areas [Hirsch and Curcio(1989)]. In order to efficiently and adaptively process visual information, human visual systems iteratively process spatial glimpses and focus on salient areas [Larochelle and Hinton(2010)].
Cross-modal attention. Attention mechanism is a widely-used technique in multi-modal settings, especially where certain modalities should be processed conditioning on other modalities. Visual question answering (VQA) task is a well-known example for such tasks. Given an image and natural language question, the task is to predict an answer such as counting the number, inferring the position or the attributes of the targets. VQA task can be seen as a set of dynamically changing tasks where the provided image should be processed according to the given question. Attention mechanism softly chooses the task(question)-relevant aspects in the image features. As suggested in [Yang et al.(2016)Yang, He, Gao, Deng, and Smola], attention maps for the image features are produced from the given question, and it act as queries to retrieve question-relevant features. The final answer is classified with the stacked images features. Another way of doing this is to use bi-directional inferring, producing attention maps for both text and images, as suggested in [Nam et al.(2017)Nam, Ha, and Kim]. In such literatures, attention maps are used as an effective way to solve tasks in a conditional fashion, but they are trained in separate stages for task-specific purposes.
Self-attention. There have been various approaches to integrate attention in DNNs, jointly training the feature extraction and attention generation in an end-to-end manner. A few attempts [Wang et al.(2017)Wang, Jiang, Qian, Yang, Li, Zhang, Wang, and Tang, Hu et al.(2017)Hu, Shen, and Sun] have been made to consider attention as an effective solution for general classification task. Wang et al\bmvaOneDothave proposed Residual Attention Networks which use a hour-glass module to generate 3D attention maps for intermediate features. Even the architecture is resistant to noisy labels due to generated attention maps, the computational/parameter overhead is large because of the heavy 3D map generation process. Hu et al\bmvaOneDothave proposed a compact ‘Squeeze-and-Excitation’ module to exploit the inter-channel relationships. Although it is not explicitly stated in the paper, it can be regarded as an attention mechanism applied upon channel axis. However, they miss the spatial axis, which is also an important factor in inferring accurate attention map.
Adaptive modules. Several previous works use adaptive modules that dynamically changes their output according to their inputs. Dynamic Filter Network [Jia et al.(2016)Jia, De Brabandere, Tuytelaars, and Gool] proposes to generate convolutional features based on the input features for flexibility. Spatial Transformer Network [Jaderberg et al.(2015b)Jaderberg, Simonyan, Zisserman, et al.] adaptively generates hyper-parameters of affine transformations using input feature so that target area feature maps are well aligned finally. This can be seen as a hard attention upon the feature maps. Deformable Convolutional Network [Dai et al.(2017)Dai, Qi, Xiong, Li, Zhang, Hu, and Wei] uses deformable convolution where pooling offsets are dynamically generated from input features, so that only the relevant features are pooled for convolutions. Similar to the above approaches, BAM is also a self-contained adaptive module that dynamically suppress or emphasize feature maps through attention mechanism.
In this work, we exploit both channel and spatial axes of attention with a simple and light-weight design. Furthermore, we find an efficient location to put our module - bottleneck of the network.
Bottleneck Attention Module
Spatial attention branch.
Combine two attention branches.
Experiments
We evaluate BAM on the standard benchmarks: CIFAR-100, ImageNet-1K for image classification and VOC 2007, MS COCO for object detection. In order to perform better apple-to-apple comparisons, we first reproduce all the reported performance of networks in the PyTorch framework [pyt()] and set as our baselines [He et al.(2016a)He, Zhang, Ren, and Sun, Zagoruyko and Komodakis(2016), Xie et al.(2016)Xie, Girshick, Dollár, Tu, and He, Huang et al.(2016a)Huang, Liu, Weinberger, and van der Maaten]. Then we perform extensive experiments to thoroughly evaluate the effectiveness of our final module. Finally, we verify that BAM outperforms all the baselines without bells and whistles, demonstrating the general applicability of BAM across different architectures as well as different tasks. Table 5, Table 6, Table 7, Table 8, Table 9 can be found at supplemental material.
The CIFAR-100 dataset [Krizhevsky and Hinton()] consists of 60,000 3232 color images drawn from 100 classes. The training and test sets contain 50,000 and 10,000 images respectively. We adopt a standard data augmentation method of random cropping with 4-pixel padding and horizontal flipping for this dataset. For pre-processing, we normalize the data using RGB mean values and standard deviations.
In Table 1, we perform an experiment to determine two major hyper-parameters in our module, which are dilation value and reduction ratio, based on the ResNet50 architecture. The dilation value determines the sizes of receptive fields in the spatial attention branch. Table 1 shows the comparison result of four different dilation values. We can clearly see the performance improvement with larger dilation values, though it is saturated at the dilation value of 4. This phenomenon can be interpreted in terms of contextual reasoning, which is widely exploited in dense prediction tasks [Yu and Koltun(2015), Long et al.(2015)Long, Shelhamer, and Darrell, Bell et al.(2016)Bell, Lawrence Zitnick, Bala, and Girshick, Chen et al.(2016)Chen, Papandreou, Kokkinos, Murphy, and Yuille, Zhu et al.(2017)Zhu, Zhao, Wang, Zhao, Wu, and Lu]. Since the sequence of dilated convolutions allows an exponential expansion of the receptive field, it enables our module to seamlessly aggregate contextual information. Note that the standard convolution (i.e. dilation value of 1) produces the lowest accuracy, demonstrating the efficacy of a context-prior for inferring the spatial attention map. The reduction ratio is directly related to the number of channels in both attention branches, which enable us to control the capacity and overhead of our module. In Table 1, we compare performance with four different reduction ratios. Interestingly, the reduction ratio of 16 achieves the best accuracy, even though the reduction ratios of 4 and 8 have higher capacity. We conjecture this result as over-fitting since the training losses converged in both cases. Based on the result in Table 1, we set the dilation value as 4 and the reduction ratio as 16 in the following experiments.
Separate or Combined branches.
In Table 1, we conduct an ablation study to validate our design choice in the module. We first remove each branch to verify the effectiveness of utilizing both channel and spatial attention branches. As shown in Table 1, although each attention branch is effective to improve performance over the baseline, we observe significant performance boosting when we use both branches jointly. This shows that combining the channel and spatial branches together play a critical role in inferring the final attention map. In fact, this design follows the similar aspect of a human visual system, which has ‘what’ (channel) and ‘where’ (spatial) pathways and both pathways contribute to process visual information [Larochelle and Hinton(2010), Chen et al.(2017)Chen, Zhang, Xiao, Nie, Shao, and Chua].
Combining methods.
We also explore three different combining strategies: element-wise maximum, element-wise product, and element-wise summation. Table 1 summarizes the comparison result for the three different implementations. We empirically confirm that element-wise summation achieves the best performance. In terms of the information flow, the element-wise summation is an effective way to integrate and secure the information from the previous layers. In the forward phase, it enables the network to use the information from two complementary branches, channel and spatial, without losing any of information. In the backward phase, the gradient is distributed equally to all of the inputs, leading to efficient training. Element-wise product, which can assign a large gradient to the small input, makes the network hard to converge, yielding the inferior performance. Element-wise maximum, which routes the gradient only to the higher input, provides a regularization effect to some extent, leading to unstable training since our module has few parameters. Note that all of three different implementations outperform the baselines, showing that utilizing each branch is crucial while the best-combining strategy further boosts performance.
Comparison with placing original convblocks.
In this experiment, we empirically verify that the significant improvement does not come from the increased depth by naively adding the extra layers to the bottlenecks. We add auxiliary convolution blocks which have the same topology with their baseline convolution blocks, then compare it with BAM in Table 1. we can obviously notice that plugging BAM not only produces superior performance but also puts less overhead than naively placing the extra layers. It implies that the improvement of BAM is not merely due to the increased depth but because of the effective feature refinement.
Bottleneck: The efficient point to place BAM.
We empirically verify that the bottlenecks of networks are the effective points to place our module BAM. Recent studies on attention mechanisms [Hu et al.(2017)Hu, Shen, and Sun, Wang et al.(2017)Wang, Jiang, Qian, Yang, Li, Zhang, Wang, and Tang] mainly focus on modifications within the ‘convolution blocks’ rather than the ‘bottlenecks’. We compare those two different locations by using various models on CIFAR-100. In Table 2, we can clearly observe that placing the module at the bottleneck is effective in terms of overhead/accuracy trade-offs. It puts much less overheads with better accuracy in most cases except PreResNet 110 [He et al.(2016b)He, Zhang, Ren, and Sun].
2 Classification Results on CIFAR-100
In Table 3, we compare the performance on CIFAR-100 after placing BAM at the bottlenecks of state-of-the-art models including [He et al.(2016a)He, Zhang, Ren, and Sun, He et al.(2016b)He, Zhang, Ren, and Sun, Zagoruyko and Komodakis(2016), Xie et al.(2016)Xie, Girshick, Dollár, Tu, and He, Huang et al.(2016a)Huang, Liu, Weinberger, and van der Maaten]. Note that, while ResNet101 and ResNeXt29 16x64d networks achieve 20.00% and 17.25% error respectively, ResNet50 with BAM and ResNeXt29 8x64d with BAM achieve 20.00% and 16.71% error respectively using only half of the parameters. It suggests that our module BAM can efficiently raise the capacity of networks with a fewer number of network parameters. Thanks to our light-weight design, the overall parameter and computational overheads are trivial.
3 Classification Results on ImageNet-1K
The ILSVRC 2012 classification dataset [Deng et al.(2009)Deng, Dong, Socher, Li, Li, and Fei-Fei] consists of 1.2 million images for training and 50,000 for validation with 1,000 object classes. We adopt the same data augmentation scheme with [He et al.(2016a)He, Zhang, Ren, and Sun, He et al.(2016b)He, Zhang, Ren, and Sun] for training and apply a single-crop evaluation with the size of 224224 at test time. Following [He et al.(2016a)He, Zhang, Ren, and Sun, He et al.(2016b)He, Zhang, Ren, and Sun, Huang et al.(2016b)Huang, Sun, Liu, Sedra, and Weinberger], we report classification errors on the validation set. ImageNet classification benchmark is one of the largest and most complex image classification benchmark, and we show the effectiveness of BAM in such a general and complex task. We use the baseline networks of ResNet [He et al.(2016a)He, Zhang, Ren, and Sun], WideResNet [Zagoruyko and Komodakis(2016)], and ResNeXt [Xie et al.(2016)Xie, Girshick, Dollár, Tu, and He] which are used for ImageNet classification task. More details are included in the supplementary material.
As shown in Table 3, the networks with BAM outperform all the baselines once again, demonstrating that BAM can generalize well on various models in the large-scale dataset. Note that the overhead of parameters and computation is negligible, which suggests that the proposed module BAM can significantly enhance the network capacity efficiently. Another notable thing is that the improved performance comes from placing only three modules overall the network. Due to space constraints, further analysis and visualizations for success and failure cases of BAM are included in the supplementary material.
4 Effectiveness of BAM with Compact Networks
The main advantage of our module is that it significantly improves performance while putting trivial overheads on the model/computational complexities. To demonstrate the advantage in more practical settings, we incorporate our module with compact networks [Howard et al.(2017)Howard, Zhu, Chen, Kalenichenko, Wang, Weyand, Andreetto, and Adam, Iandola et al.(2016)Iandola, Han, Moskewicz, Ashraf, Dally, and Keutzer], which have tight resource constraints. Compact networks are designed for mobile and embedded systems, so the design options have computational and parametric limitations.
As shown in Table 3, BAM boosts the accuracy of all the models with little overheads. Since we do not adopt any squeezing operation [Howard et al.(2017)Howard, Zhu, Chen, Kalenichenko, Wang, Weyand, Andreetto, and Adam, Iandola et al.(2016)Iandola, Han, Moskewicz, Ashraf, Dally, and Keutzer] on our module, we believe there is more room to be improved in terms of efficiency.
5 MS COCO Object Detection
We conduct object detection on the Microsoft COCO dataset [Lin et al.(2014)Lin, Maire, Belongie, Hays, Perona, Ramanan, Dollár, and Zitnick]. According to [Bell et al.(2016)Bell, Lawrence Zitnick, Bala, and Girshick, Liu et al.(2016)Liu, Anguelov, Erhan, Szegedy, Reed, Fu, and Berg], we trained our model using all the training images as well as a subset of validation images, holding out 5,000 examples for validation. We adopt Faster-RCNN [Ren et al.(2015)Ren, He, Girshick, and Sun] as our detection method and ImageNet pre-trained ResNet101 [He et al.(2016a)He, Zhang, Ren, and Sun] as a baseline network. Here we are interested in improving performance by plugging BAM to the baseline. Because we use the same detection method of both models, the gains can only be attributed to our module BAM. As shown in the Table 4, we observe significant improvements from the baseline, demonstrating generalization performance of BAM on other recognition tasks.
6 VOC 2007 Object Detection
We further experiment BAM on the PASCAL VOC 2007 detection task. In this experiment, we apply BAM to the detectors. We adopt the StairNet [Sanghyun et al.(2018)Sanghyun, Soonmin, and So] framework, which is one of the strongest multi-scale method based on the SSD [Liu et al.(2016)Liu, Anguelov, Erhan, Szegedy, Reed, Fu, and Berg]. We place BAM right before every classifier, refining the final features before the prediction, enforcing model to adaptively select only the meaningful features. The experimental results are summarized in Table 4. We can clearly see that BAM improves the accuracy of all strong baselines with two backbone networks. Note that accuracy improvement of BAM comes with a negligible parameter overhead, indicating that enhancement is not due to a naive capacity-increment but because of our effective feature refinement. In addition, the result using the light-weight backbone network [Howard et al.(2017)Howard, Zhu, Chen, Kalenichenko, Wang, Weyand, Andreetto, and Adam] again shows that BAM can be an interesting method to low-end devices.
7 Comparison with Squeeze-and-Excitation[Hu et al.(2017)Hu, Shen, and Sun]
We conduct additional experiments to compare our method with SE in CIFAR-100 classification task. Table 5 summarizes all the results showing that BAM outperforms SE in most cases with fewer parameters. Our module requires slightly more GFLOPS but has much less parameters than SE, as we place our module only at the bottlenecks not every conv blocks.
Conclusion
We have presented the bottleneck attention module (BAM), a new approach to enhancing the representation power of a network. Our module learns what and where to focus or suppress efficiently through two separate pathways and refines intermediate features effectively. Inspired by a human visual system, we suggest placing an attention module at the bottleneck of a network which is the most critical points of information flow. To verify its efficacy, we conducted extensive experiments with various state-of-the-art models and confirmed that BAM outperforms all the baselines on three different benchmark datasets: CIFAR-100, ImageNet-1K, VOC2007, and MS COCO. In addition, we visualize how the module acts on the intermediate feature maps to get a clearer understanding. Interestingly, we observed hierarchical reasoning process which is similar to human perception procedure. We believe our findings of adaptive feature refinement at the bottleneck is helpful to the other vision tasks as well.