Comprehensive Attention Self-Distillation for Weakly-Supervised Object Detection

Zeyi Huang, Yang Zou, Vijayakumar Bhagavatula, Dong Huang

Introduction

Visual object detection has achieved remarkable progress in the last decade thanks to the advances of Convolutional Neural Networks (CNNs) . An integral part of the achievement is the availability of large-scale training data with precise bounding-box annotations (PASCAL VOC , MS-COCO , etc). However, obtaining such fine-grained annotations at a large scale is labor-intensive and time-consuming, which drove many researchers to explore the weakly-supervised setting. Weakly-Supervised Object Detection (WSOD) aims to learn object detectors with only the image-level category labels indicating whether an image contains an object or not.

Most previous methods for WSOD are based on the Multiple Instance Learning (MIL) . These methods regard images as bags and object proposals as instances. A positive bag contains at least one positive instance while all instances being negative in a negative bag. WSOD instance classifiers (object detectors) are trained over these bags. Recently, leveraging the powerful representation learning capacity of CNNs, several researchers proposed end-to-end MIL networks (OICR , PCL , MIST , ) with promising WSOD performances. These CNN methods regard the instance classification (object detection) problem as a latent model learning within a bag classification (image classification) problem, where the final image scores are the aggregation of the instance scores. However, due to the under-determined and ill-posed nature of WSOD, there is still a large performance gap between the weakly-supervised detectors and fully-supervised detectors.

The existing methods have two main sets of issues as demonstrated in Fig. 1. First, in the “Biased WSOD” column of Fig. 1 (a) , there are three typical problems. Missing instance: Salient objects are easily detected while inconspicuous instances tend to be ignored. Clustered instances: multiple adjacent instances of the same category may be detected in a single bounding box. Part domination: The bounding boxes are prone to focus on the most discriminative object parts instead of the entire objects. Second, in the “Inconsistent WSOD” column of Fig. 1 (b), the same image and its different image transformations, i.e., “Original Image”, “Flipped Image” and “Scaled Image”, do not produce the same object bounding boxes.

WSOD conducts classification on object proposals (e.g., bounding boxes generated by selective search ) with image-level class labels. The object proposals receive high classification scores are considered as objects detected by WSOD. As we dive deep into the above issues from a feature learning perspective, we overlay the attention maps of object proposals that get high confidences in WSOD (Fig. 1). High intensity in attention maps corresponds to highly discrimiative and biased features learned by the WSOD networks. We observe the drawbacks of WSOD detection are closely associated with the issues in feature learning. For “Biased WSOD”, it is clear that salient objects, clustered objects, and certain object parts contain spatial features that dominate the WSOD classification. From a statistical machine learning point-of-view, feature domination is typically established by the biased feature distribution in training data. For “Inconsistent WSOD”, the different transformations of the same image are typically generated by data augmentation and are used to train the WSOD networks in different training iterations. The same class image-level labels of transformed images do not enforce spatially consistent feature learning and may lead to part domination and missing instances. Note that the inconsistency on feature localization was not an issue for full-supervised setting where augmented training data with precise bounding box labels can naturally encourage consistency.

The above observations inspire us to address WSOD issues using an attention-based feature learning method. We propose a Comprehensive Attention Self-Distillation (CASD) approach for WSOD training. To balance feature learning among objects, CASD computes the comprehensive attention aggregated from multiple transformations and feature layers of the same images. The “CASD (ours)” column of Fig. 1 (a) demonstrates that CASD generates balanced attention on less salient objects, individual objects, and entire objects, which enables WSOD detection on these objects. To enforce consistent spatial supervision on objects, CASD conducts self-distillation on the WSOD network itself, such that the comprehensive attention is approximated simultaneously by multiple transformations and layers of the same images. The “CASD (ours)” column of Fig. 1 (b) demonstrates that CASD generates consistent attention on different transformed variants of the same image, leading to consistent WSOD detection in different transformations.

By computing the comprehensive attention maps, CASD aggregates “free” resources of spatial supervision for WSOD, including image transformations and low-to-high feature layers. By conducting self-distillation on the WSOD network with the comprehensive attention maps, CASD enforces instance-balanced and spatially-consistent supervision, therefore robust bounding box localization for WSOD. CASD achieves the state-of-the-art on several standard benchmarks, e.g. PASCAL VOC 2007/2012 and MS-COCO, outperforming other methods by clear margins. Systematic ablation studies are also conducted on the effects of transformations and feature layers on CASD.

Related Work

Recent WSOD performance are significantly boosted by incorporating Multiple Instance Learning (MIL) in Convolutional Neural Networks (CNN). introduces the first end-to-end Weakly Supervised Deep Detection Network (WSDDN) with MIL, which inspired many following works. combines WSDDN and multi-stage instance classifiers into an Online Instance Classifier Refinement (OICR) framework. further improves OICR with a robust proposal generation module based on proposal clustering, namely Proposal Cluster Learning (PCL). introduces Continuation Multiple Instance Learning (C-MIL) by relaxing the original MIL loss function with a set of smoothed loss functions preventing detectors to be part dominating. proposes a multiple instance self-training framework with an online regression branch. and leverage segmentation maps to generate instance proposals with rich contextual information. and introduce detection-segmentation cyclic collaborative frameworks. Different from the above methods that regularize the WSOD outputs, CASD directly enforces comprehensive and consistent WSOD feature learning.

2 Attention Mechanism in Computer Vision DNNs

The attention mechanism provides a fine-grained view of the features learned in CNNs. introduce the connection between class-wise attention maps and image-level class labels. Due to its explicit spatial clues, attention maps have been used to improve computer vision tasks in two ways: (1) Re-weighting features. re-weight features with spatial-wise attention maps. improve supervised tasks by inverting gradient-based spatial-wise and channel-wise attention. (2) Loss regularization. introduces a consistency loss between attention maps under different input transformations for image classification. propose cross-layer consistency losses over attention maps for image classification and lane detection. To the best of our knowledge, CASD is the first attempt to explore attention regularization for WSOD. Moreover, existing methods for other tasks only encourage consistency of features, rather than the completeness of features. Specifically, the features tend to focus on object parts but fail to localize less salient objects in WSOD. CASD encourages both consistent and spatially complete feature learning guided by the comprehensive attention maps, which explicitly addresses WSOD issues above.

3 Knowledge Distillation

Knowledge distillation is a CNN training process to transfer knowledge from teacher networks to student networks. It has found wide applications in model compression, incremental learning, and continual learning. The student networks mimic the the teacher networks on predictive probabilities , intermediate features , or attention maps of intermediate neural activations . In contrast, CASD is a knowledge self-distillation process that transfers the comprehensive attention knowledge within the WSOD model itself, rather than another teacher model, across multiple views of the same data.

Method

Proposal Feature Extractor. From an input image I\mathbf{I}, the object proposals R={R1,R2,...,RN}\mathbf{R}=\{R_{1},R_{2},...,R_{N}\} are generated by Selective Search where NN is total number of proposals (bounding boxes). Then, a CNN backbone is used to extract the image feature maps for I\mathbf{I}, from which the proposal feature maps are extracted respectively. Lastly, these proposal feature maps are fed to a Region-of-Interest (RoI) Pooling layer and two Fully-Connected (FC) layers to obtain proposal feature vectors.

Multiple Instance Learning (MIL) Head. The MIL head learns the latent instance classifiers (detectors). This module takes the above-obtained proposal feature vectors as input and conducts the image-level MIL classification, regarding detection as a latent model learning problem.

2 Comprehensive Attention Distillation

Building upon , Comprehensive Attention Self-Distillation (CASD) encourages consistent and balanced representation learning in two folds: Input-wise CASD and Layer-wise CASD.

where Ar(m,n)\mathbf{A}_{r}(m,n) denotes the proposal attention magnitude at spatial location (m,n)(m,n) of proposal RrR_{r}. S(⋅)S(\cdot) is the element-wise Sigmoid operator. The proposal attention maps Arflip\mathbf{A}_{r}^{flip} and Arscale\mathbf{A}_{r}^{scale} from Fflip\mathbf{F}^{flip} and Fscale\mathbf{F}^{scale} can be computed in the same way as Eq. 1.

The example in Fig. 3 (a) shows that attention maps Ar\mathbf{A}_{r}, Arflip\mathbf{A}^{flip}_{r} and Arscale\mathbf{A}^{scale}_{r} focus on different parts of proposal RrR_{r}. This is the common case in WSOD: for input image under different transformations, the feature distribution within class does not correlate with the feature distribution between classes. Unless regularized, the image-level classification of WSOD always relies on the most discriminative features, which only correspond to the proposals on salient objects or object parts.

The union of these attention maps (denoted by ArIW\mathbf{A}^{IW}_{r}) always covers more comprehensive parts of the object than individual attention maps,

where max(⋅)max(\cdot) is the element-wise max operator. We define the input-wise CASD based on ArIW\mathbf{A}^{IW}_{r} as a refinement problem on the WSOD network: updating the WSOD feature extractor such that any transformed variants of the same image should generate comprehensive attention close to ArIW\mathbf{A}^{IW}_{r}. To this end, ArIW\mathbf{A}^{IW}_{r} are simultaneously approximated from individual attention maps. For each kthk_{th} refinement branch , the IW-CASD loss function is defined as

NKN_{K} is the total number of selected proposals in the kthk_{th} branch. The computation of ArIWA^{IW}_{r} and updating of the WSOD feature extractor are alternative steps in training. We consider this is a knowledge distillation process from ArIWA^{IW}_{r} to the WSOD network itself, and hereby name it as self-distillation.

However, the alternating process above may also lead to local optimum. In particular, the individual attention map containing more high-intensity elements may dominate attention maps associated with other transformations. We balance the distillation among different transformations by independently applying Inverted Attention (IA) on each individual attention. For each attention, IA randomly masks out top features highlighted by attention maps and force more features to be activated. This training technique regularizes CASD and produces better results (see ablation study in Section 4).

Additionally, the aggregated proposal score matrices are used in the WSOD network training for each transformed variant of the same image. Specifically, in the MIL head, IW-CASD takes several transformed variants of the same image to generate groups of proposal instance score matrices for each branch (K+1K+1 in total). Within each kthk_{th} group, each score matrix provides a transformed view of the original image for detection. Thus it is natural for WSOD to aggregate the score matrices within each kthk_{th} group to form a robust proposal score matrix as xˉk=13(xk+xflipk+xscalek)\bar{\mathbf{x}}^{k}=\frac{1}{3}(\mathbf{x}^{k}+\mathbf{x}_{flip}^{k}+\mathbf{x}_{scale}^{k}).

Layer-wise(LW) CASD (LW-CASD) operates on proposal feature maps of image I\mathbf{I} produced at multiple layers of WSOD feature extractor (see Fig. 3 (b)). The feature extractor network consists of QQ convolutional blocks B1,...,BQB_{1},...,B_{Q}, each of which outputs a feature map FB1,...,FBQ\mathbf{F}^{B_{1}},...,\mathbf{F}^{B_{Q}}, respectively. FBQ\mathbf{F}^{B_{Q}} is used as the feature map for generating proposal feature vectors in the MIL head. As illustrated by the example in Fig. 3 (b), at the early layers of a CNN, the network focuses on local low-level features, while at deeper layers, the network tends to focus on global semantic features. Conducting CASD over layers enriches the proposal features with more granularity.

In LW-CASD, we use RoI pooling at different feature blocks to generate proposal feature maps, such that they share the same spatial size. From these proposal attention maps, the layer-wise attention maps ABq\mathbf{A}^{B_{q}}s are generated in a similar way as Eq. 1, and aggregated to obtain the comprehensive attention maps,

where max(⋅)max(\cdot) is the element-wise max operator. For the kthk_{th} refinement branch, the IW-CASD loss is

The choice of layers in IW-CASD are studied in Section 4.

Remarks. 1) Channel-wise average pooling+sigmoid in Eq. 1 is not the only choice for attention estimation. We also explored other mechanisms such as Grad-CAM , but there is only a negligible performance difference. Eq. 1 is used due to its simplicity and computation efficiency. 2) Besides flipping and scaling, any other image transformations could also be used in IW-CASD. 3) In Eq. 3 and Eq. 5, gradients are not back-propagated to the comprehensive attention maps ArIWA_{r}^{IW} or ArLWA_{r}^{LW}.

Overall Loss Function. The overall loss for training a CASD-based WSOD network is composed of the losses from Sec. 3.1 and Sec. 3.2.

where α\alpha, β\beta and γ\gamma balance the weights of different losses. KK is the number of refinement branches.

Experimental Results

Datasets and Metrics. Three standard WSOD benchmarks, PASCAL VOC 2007, VOC 2012 and MS-COCO , are used in our experiments. Both VOC 2007 and VOC 2012 contain 20 object classes with an additional background class. In VOC 2007, the total of 9,9629,962 images are split into three subsets: 2,5012,501 for training, 2,5102,510 for validation and 4,9514,951 testing. In VOC 2012, all the 22,53122,531 images are split into the 5,7175,717 training images, 5,8235,823 validation images, the rest 10,99110,991 test images. For both datasets, we followed the standard routine in WSOD to train on the train+val set and evaluate on the test set. For MS-COCO trainval set, the train set (82,78382,783 images) is used for training and the val set (40K images) is used for testing. Only image-level labels are utilized in training. For evaluation, a predicted bounding box is considered to be positive if it has an IoU>0.5>0.5 with the ground-truth. For VOC, mean Average Precision mAP0.5 (IoU threshold at 0.50.5) is reported. For MS-COCO, we report mAP0.5 and mAP (averaged over IoU thresholds in [.5:0.05:.95][.5:0.05:.95]).

Implementation details. All experiments were implemented in PyTorch. The VGG16 and ResNet50 pre-trained on ImageNet are used as WSOD backbones. Batch size is set to be TT that is the number of input transformations. The maximum iteration numbers are set to be 80K80K, 160K160K and 200K200K for VOC 2007, VOC 2012, and MS-COCO respectively. The whole WSOD network is optimized in an end-to-end way by stochastic gradient descent (SGD) with a momentum of 0.90.9, an initial learning rate of 0.0010.001 and a weight decay of 0.00050.0005. The learning rate will decay with a factor of 1010 at the 4040kth, 8080kth, and 120120kth iteration for VOC 2007, VOC 2012 and MS-COCO, respectively. The total number of refinement branches KK is set to be 22. The confidence threshold for Non-Maximum Suppression (NMS) is 0.30.3. For all experiments, we set α=0.1\alpha=0.1, β=0.05\beta=0.05 and γ=0.1\gamma=0.1 in the total loss. Note that we reimplemented the OICR loss in Pytorch and found that using a fixed α=0.1\alpha=0.1 is slightlty better than the adaptive weighting policy in vanilla OICR (48.9%48.9\% mAP0.5 vs. 48.3%48.3\% mAP0.5 on VOC 2007). Thus we keep the fixed weight α\alpha for the OICR loss. Selective Search is used to generate about 2,0002,000 object proposals for each image.

To have a fair comparison with other methods, following , multi-level scaling and horizontal flipping data augmentation are conducted in training. Multi-level scaling is also used in testing. Specifically, in the data augmentation, the short edges of input images are randomly re-scaled to a scale in {480,576,688,864,1200}\{480,576,688,864,1200\}, and the longest image edges are capped to 2,0002,000, then a random horizontal flipping is randomly conducted on the scaled images. In evaluation, input images are augmented with all five scales.

We conducted three sets of ablation studies on the VOC 2007 with metric mAP0.5 (%\%) in Table 3- 3. All results are based on the VGG16 backbone.

CASD main configurations. We conducted ablation studies on the main components of CASD in Table 3 under the mAP0.5 metric. Our baseline achieves 48.9%48.9\% mAP0.5. With the proposed Input-wise CASD (“+IW”), the baseline is boosted to 54.1%54.1\% mAP0.5 while the proposed Layer-wise CASD (+LW) improves the baseline to 52.3%52.3\% mAP0.5. Both IW-CASD and LW-CASD can consistently improve the baseline with a clear 5.2%5.2\% and 3.4%3.4\% gain respectively. Combining IW with LW (“+IW+LW”) can further boost the performance to 55.3%55.3\% mAP0.5 that is 6.4%6.4\% better than the baseline. In addition, IW-CASD without Inverted Attention (IA) and proposal score aggregation (PSA), “+IW w/o IA+PSA”, could achieve 51.4%51.4\% mAP0.5 which is 2.5%2.5\% superior to the baseline. Proposal score aggregation brings a 1.2%1.2\% mAP0.5 performance gain, leading to 52.6%52.6\% mAP0.5 in IW-CASD without Inverted Attention (“+IW w/o IA”). And Inverted Attention (+IW) has a 1.5%1.5\% mAP0.5 performance boost, justifying the effectiveness of IA.

Then we further demonstrate the validity of regression branch and stronger augmentation. With the regression branch, “+IW+LW+Reg” achieves 56.1%56.1\% mAP0.5 that is 0.8%0.8\% better than “+IW+LW”. Moreover, with stronger data augmentation of “flip+scale+color”, we can achieve the best CASD performance (“+IW+LW+Reg∗”) 56.8%56.8\% mAP0.5. Here the “Color” augmentation applies some photo-metric distortions to the images which is the same as those described in .

CASD layer configurations. Built upon the baseline, additional ablation study on layer configurations of LW-CASD are shown in Table 3. The BiB_{i} is the ithi_{th} convolutional block of VGG16 and B4B_{4} denotes the last block before the FC layer for image classification. As shown in the table, the best result 52.3%52.3\% mAP0.5 is obtained with LW-CASD using B4+B3+B2B_{4}+B_{3}+B_{2} blocks. As demonstrated in Fig. 3 (b), these results indicate that middle-level feature attention maps (e.g. B2B_{2} and B3B_{3}) encode balanced discriminative clues (more spatially distributed than B1B_{1}) for WSOD, while the low-level feature attention maps (e.g. B1B_{1}) contain noisy spatial information which may deteriorate the performance. Besides Table 3, all other experiments in this work use the B4+B3+B2B_{4}+B_{3}+B_{2} configuration in CASD.

Attention regularization strategies. Besides our attention self-distillation, there are other regularization strategies emerged for semi-supervised object detection and multi-label classification such as Prediction Consistency and Attention Consistency . For a thorough demonstration of the advantage of our attention distillation strategy, we implement these strategies under the WSOD setting and compare them with the CASD. Similar to , prediction consistency in WSOD is implemented by minimizing the JS-Divergence between predictions of differently transformed data. Attention consistency in WSOD force the attention of the last proposal feature maps to be consistent under different input transformations based on the MSE loss.

We compared IW-CASD (“+IW w/o IA”) with the other two methods in Table 3. IW-CASD is superior to both prediction and attention consistency strategies by a 2.4%2.4\% mAP0.5 and 1.6%1.6\% mAP0.5 performance gain respectively. This validates the superiority of CASD over other attention consistency methods.

Sensitivity analysis on loss weight γ\gamma: For the loss weight γ\gamma in Eq. 6, we evaluated CASD on VOC 2007 with γ=0.05,0.075,0.1,0.15,0.2\gamma=0.05,0.075,0.1,0.15,0.2, and got mAP 54.1%,54.7%,55.3%,55.0%,55.0%54.1\%,54.7\%,55.3\%,55.0\%,55.0\% respectively, which demonstrates that CASD is robust to γ\gamma.

2 Comparison with the State-of-the-arts

Here we experimentally compare the full version of CASD with other state-of-the-art methods. In Table 5, with the same VGG16 backbone, CASD reaches the new state-of-the-art mAP0.5 of 56.8%56.8\% and 53.6%53.6\% on VOC 2007 and VOC 2012, which are 1.9%1.9\% and 1.5%1.5\% higher than the latest state-of-the-art (MIST(+Reg.) ). In the results on MS-COCO (Table 5). With the VGG16 backbone, CASD produces 12.8%12.8\% mAP and 26.4%26.4\% mAP0.5, outperforming the VGG16 version of MIST(+Reg.) by clear margins of 1.4%1.4\% and 2.1%2.1\%. With the ResNet50 backbone, CASD achieves the state-of-the-art 13.9%13.9\% mAP and 27.8%27.8\% mAP0.5 outperforming MIST(+Reg.) by clear margins of 1.3%1.3\% mAP and 1.7%1.7\% mAP0.5. Qualitative visualization of detection results can be found in Fig. 4 and 5 in the Appendix.

Conclusion

In this paper, we proposed the Comprehensive Attention Self-Distillation (CASD) algorithm to regularize the WSOD training. CASD aggregates “free” resources of spatial supervision within the WSOD network, such as the different attentions produced under multiple image transformations and low-to-high feature layers. Through self-distillation on the WSOD network, CASD enforces instance-balanced and spatially-consistent supervision over objects and achieves the new state-of-the-art WSOD results. As a training module, we believe CASD can be generalized and benefit other weakly-supervised and semi-supervised tasks, such as instance segmentation, pose estimation.

Broader Impact

This paper pushes the frontier of the weakly supervised object detection and reduce its performance gap with the supervised detection. This work is also a general regularization approach that may benefit semi-supervised learning, weakly-supervised learning, and self-supervised representation learning. The ultimate research vision is to potentially relieve the burden of human annotations on training data. This effort may reduce the cost, balance human bias, accelerate the evolution of machine perception technology, and help us to understand how to enable learning with minimal supervision.

Acknowledgments and Disclosure of Funding

This work was supported by the Intelligence Advanced Research Projects Activity (IARPA) via Department of Interior/ Interior Business Center (DOI/IBC) contract number D17PC00340.

References

Appendix

Appendix A More Experimental Results

We first present qualitative results of CASD-WSOD in Fig. 4. Recall that WSOD conducts classification on object proposals (e.g., bounding boxes generated by Selective Search ) with image-level class labels. The object proposals receive high classification scores are considered as objects detected by WSOD. In Fig. 4, only the attention maps of high confidence proposals by WSOD detection are overlaid on input images.

Fig. 4 shows both the success and the failure cases of CASD. When the objects are relatively big and are completely visible, CASD tends to succeed. The failure cases mainly occur on objects that are small or under heavy occlusion. We deduce that there are two main factors contributing to the failure: (1) The Selective Search (SS) algorithm may not generate good proposals for heavily occluded objects. This could be improved by using better objectness proposal generators such as the RPN of Faster-RCNN. (2) The incomplete appearance of the occluded objects make CASD difficult to learn the long-range dependency among the object parts. This could be improved by hard-sample mining in CASD training.

Fig. 5 compares results of MIST and CASD. CASD has better success in detecting high-quality bounding boxes than MIST. This localization advantages of CASD benefit from its learning of comprehensive attention (see the bottom row of Fig. 5). We further demonstrate the localization quality of CASD in Table 6.

A.2 CorLoc on Trainval Sets

Correct Localization (CorLoc) was used in some previous works to evaluate performance on the VOC 2007 and VOC 2012 trainval sets. CorLoc only evaluates the localization accuracy of detectors. In all those works, CorLoc was reported when WSOD is trained on both the training and validation sets, and tested on both sets. This is probably why CorLoc was mostly used for ablation, not for the main comparison between algorithms.

For completeness, we provide this additional metric in Table 6. CASD has the best overall localization accuracy among all compared methods.

A.3 More Ablation Studies and Clarification

Different γ\gammas for LIWL_{IW} and LLWL_{LW}: In Eq. 6, we have a single loss weight γ\gamma for both the IW-CASD loss LIWk\mathcal{L}_{IW}^{k} and LW-CASD loss LLWk\mathcal{L}_{LW}^{k}. Another policy is to set different loss weights for the two losses respectively. Which way is better? So we conduct the following ablation study on VOC 2007. Fixing γIW=0.1\gamma_{IW}=0.1, CASD gets 55.0%,55.5%,55.3%,54.8%,54.8%55.0\%,55.5\%,55.3\%,54.8\%,54.8\% mAP0.5 when γLW=0.05,0.075,0.1,0.15,0.2\gamma_{LW}=0.05,0.075,0.1,0.15,0.2 respectively. Fixing γLW=0.1\gamma_{LW}=0.1, CASD gets 54.3%,54.6%,55.3%,55.1%,54.9%54.3\%,54.6\%,55.3\%,55.1\%,54.9\% mAP0.5 when γIW=0.05,0.075,0.1,0.15,0.2\gamma_{IW}=0.05,0.075,0.1,0.15,0.2 respectively. At LIW=0.1L_{IW}=0.1 and LLW=0.075L_{LW}=0.075, CASD achieves 55.5%55.5\% mAP0.5 which is only 0.2%0.2\% better than 55.3%55.3\% mAP0.5 (the best performance of single γ\gamma). Thus we conclude that single γ\gamma is a good trade-off between performance and hyper-parameter tuning.

Evidence for consistency and completeness in CASD: First, the attention maps and predicted bounding boxes in Fig. 1 and Fig. 5 compare the results of OICR/MIST and CASD, qualitatively demonstrating CASD gets more consistent and complete object features. Second, on the horizontal flipped VOC 2007 test set, CASD achieves 56.5%56.5\% mAP0.5 which is similar to 56.8%56.8\% of the unflipped test set. This indicates that CASD is consistent w.r.t flipping.

CASD with Grad-CAM: In CASD, we utilize channel-wise average pooling++sigmoid to get the attention map. We conduct an ablation study on layer-wise CASD with Grad-CAM that gets 52.0%52.0\% mAP0.5 on VOC 2007. This is slightly worse than 52.6%52.6\% mAP0.5 of layer-wise CASD with channel-wise average pooling++sigmoid. Thus the latter is computationally more efficient and adopted in our paper.

Clarification on the branch number of kk in OICR: The vanilla OICR has 3 OICR branches (k=3)(k=3), and suggests the larger kk the better results. We only use k=2k=2 in all our experiments due to our GPU limitations. Thus CASD may achieve better results by setting k=3k=3 on GPU with sufficient memory.