Instance-aware, Context-focused, and Memory-efficient Weakly Supervised Object Detection

Zhongzheng Ren, Zhiding Yu, Xiaodong Yang, Ming-Yu Liu, Yong Jae Lee, Alexander G. Schwing, Jan Kautz

Introduction

Recent works on object detection have achieved impressive results. However, the training process often requires strong supervision in terms of precise bounding boxes. Obtaining such annotations at a large scale can be costly, time-consuming, or even infeasible. This motivates weakly supervised object detection (WSOD) methods where detectors are trained with weaker forms of supervision such as image-level category labels. These works typically formulate WSOD as a multiple instance learning task, treating the set of object proposals in each image as a bag. The selection of proposals that truly cover objects is modeled using learnable latent variables.

While alleviating the need for precise annotations, existing weakly supervised object detection methods often face three major challenges due to the under-determined and ill-posed nature, as demonstrated in Fig. 1:

(1) Instance Ambiguity. This arguably the biggest challenge which subsumes two common types of issues: (a) Missing Instances: Less salient objects in the background with rare poses and smaller scales are often ignored (top row in Fig. 1). (b) Grouped Instances: Multiple instances of the same category are grouped into a single bounding box when spatially adjacent (middle row in Fig. 1). Both issues are caused by bigger or more salient boxes receiving higher scores than smaller or less salient ones.

(2) Part Domination. Predictions tend to be dominated by the most discriminative parts of an object (Fig. 1 bottom). This issue is particularly pronounced for classes with big intra-class difference. For example, on classes such as animals and people, the model often turns into a ‘face detector’ as faces are the most consistent appearance signal.

(3) Memory Consumption. Existing proposal generation methods often produce dense proposals. Without ground-truth localization, maintaining a large number of proposals is necessary to achieve a reasonable recall rate and good performance. This requires a lot of memory, especially for video object detection. Due to the large number of proposals, most memory is consumed in the intermediate layers after ROI-Pooling.

To address the above three challenges, we propose a unified weakly supervised learning framework that is instance-aware and context-focused. The proposed method tackles Instance Ambiguity by introducing an advanced self-training algorithm where instance-level pseudo ground-truth, in forms of category labels and regression targets are computed by considering more instance-associative spatial diversification constraints (Sec. 4.1). The proposed method also addresses Part Domination by introducing a parametric spatial dropout termed ‘Concrete DropBlock.’ This module is learned end-to-end to adversarially maximize the detection objective, thus encouraging the whole framework to consider context rather than focusing on the most discriminative parts (Sec. 4.2). Finally, to alleviate the issue of Memory Consumption, our method adopts a sequential batch back-propagation algorithm which processes data in batches at the most memory-heavy stage. This permits the assess to larger deep models such as ResNet in WSOD, as well as the exploration of weakly supervised video object detection (Sec. 4.3).

Tackling the aforementioned three challenges via our proposed framework leads to state-of-the-art performance on several popular datasets, including COCO , VOC 2007 and 2012 . The effectiveness and robustness of each proposed module is demonstrated in detailed ablation studies, and further verified through qualitative results. Finally, we conduct additional experiments on videos and give the first benchmark for weakly supervised video object detection on ImageNet VID .

Related work

Object detection is one of the most fundamental problems in computer vision. Recent supervised methods have shown great performance in terms of both accuracy and speed. For WSOD, most methods formulate a multiple instance learning problem where input images contain a bag of instances (object proposals). The model is trained with a classification loss to select the most confident positive proposals. Modifications w.r.t. initialization , regularization , and representations have been shown to improve results. For instance, Bilen and Vedaldi proposed an end-to-end trainable architecture for this task. Follow-up works further improve by leveraging spatial relations , better optimization , and multitasking with weakly supervised segmentation .

Self-training for WSOD.

Among the above directions, self-training has been demonstrated to be seminal. Self-training uses instance-level pseudo labels to augment training and can be implemented in an offline manner : a WSOD model is first trained using any of the methods discussed above; then the confident predictions are used as pseudo-labels to train a final supervised detector. This iterative knowledge distillation procedure is beneficial since the additional supervised models learn form less noisy data and usually have better architectures for which training is time-consuming. A number of works studied end-to-end implementations of self-training: WSOD models compute and use pseudo labels simultaneously during training, which is commonly referred to as an online solution. However, these methods typically only consider the most confident predictions for pseudo-labels. Hence they tend to have overfitting issues with difficult parts and instances ignored.

Spatial dropout.

To address the above issue, an effective regularization strategy is to drop parts of spatial feature maps during training. Variants of spatial-dropout have been widely designed for supervised tasks such as classification , object detection , and human joints localization . Similar approaches have also been applied in weakly supervised tasks for better localization in detection and semantic segmentation . However, these methods are non-parametric and cannot adapt to different datasets in a data-driven manner. As a further improvement, Kingma et al. designed variational dropout where the dropout rates are learned during training. Wang et al. proposed a parametric but non-differentiable spatial-dropout trained with REINFORCE . In contrast, the proposed ‘Concrete DropBlock’ module has a parametric and differentiable structured novel form.

Memory efficient back-propagation.

Memory has always been a concern since deeper models and larger batch size often tend to yield better results. One way to alleviate this concern is to trade computation time for memory consumption by modifying the back-propagation (BP) algorithm . A suitable technique is to not store some intermediate deep net representations during forward-propagation. One can recover those by injecting small forward passes during back-propagation. Hence, the one-stage back-propagation is divided into several step-wise processes. However, this method cannot be directly applied to our model where a few intermediate layers consume most of the memory. To address it, we suggest a batch operation for the memory-heavy intermediate layers.

Background

The final score sw(c,r)s_{w}(c,r) for assigning category cc to region rr is computed via an element-wise product: sw(c,r)=sw(c∣r)sw(r∣c)∈s_{w}(c,r)=s_{w}(c|r)s_{w}(r|c)\in. During training, sw(c,r)s_{w}(c,r) is summed for all regions r∈Rr\in R to obtain the image evidence ϕw(c)=∑r∈Rsw(c,r)\phi_{w}(c)=\sum_{r\in R}s_{w}(c,r). The loss is then computed via:

where y(c)∈{0,1}y(c)\in\{0,1\} is the ground truth (GT) class label indicating image-level existence of category cc. For inference, sw(c,r)s_{w}(c,r) is used for prediction followed by standard non-maximum suppression (NMS) and thresholding.

To integrate online self-training, the region score sw(c,r)s_{w}(c,r) is often used as teacher to generate instance-level pseudo category label y^(c,r)∈{0,1}\hat{y}(c,r)\in\{0,1\} for every region r∈Rr\in R . This is done by treating the top-scoring region and its highly-overlapped neighbors as the positive examples for class cc. The extra student layer is then trained for region classification via:

where s^w(c∣r)\hat{s}_{w}(c|r) is the output of this layer. During testing, the student prediction s^w(c∣r)\hat{s}_{w}(c|r) will be used rather than sw(c,r)s_{w}(c,r). We build upon this formulation and develop two additional novel modules as described subsequently.

Approach

Image-level labels are an effective form of supervision to mine for common patterns across images. Yet inexact supervision often causes localization ambiguity. To address the mentioned three challenges caused by this ambiguity, we develop the instance-aware and context-focused framework outlined in Fig. 2. It contains a novel online self-training algorithm with ROI regression to reduce instance ambiguity and better leverage the self-training supervision (Sec. 4.1). It also reduces part-domination for classes with large intra-class variance via a novel end-to-end learnable ‘Concrete DropBlock’ (Sec. 4.2), and it is more memory friendly (Sec. 4.3).

With online or offline generated pseudo-labels , self-training helps to eliminate localization ambiguities, benefiting mainly from two aspects: (1) Pseudo-labels permit to model proposal-level supervision and inter-proposal relations; (2) Self-training can be broadly regarded as a teacher-student distillation process which has been found helpful to improve the student’s representation. We take the following dimensions into account when designing our framework:

Instance-associative: Object detection is often ‘instance-associative’: highly overlapping proposals should be assigned similar labels. Most self-training methods for WSOD ignore this and instead treat proposals independently. Instead, we impose explicit instance-associative constraints into pseudo box generation.

Representativeness: The score of each proposal in general is a good proxy for its representativeness. It is not perfect, especially in the beginning there is a tendency to focus on object parts. However, the score provides a high recall for being at least located on correct objects.

Spatial-diversity: Imposing spatial diversity to the selected pseudo-labels can be a useful self-training inductive bias. It promotes better coverage on difficult (e.g., rare appearance, poses, or occluded) objects, and higher recall for multiple instances (e.g., diverse scales and sizes).

The above constraints and criteria motivate a novel algorithm to generate diverse yet representative pseudo boxes which are instance-associative. The details are provided in Alg. 1. Specifically, we first sort all the scores across the set RR for each class cc that appears in the category-label. We then pick the top pp percent of the ranked regions to form an initial candidate pool R′(c)R^{\prime}(c). Note that the size of the candidate pool R′(c)R^{\prime}(c), i.e., ∣R′(c)∣|R^{\prime}(c)| is image-adaptive and content-dependent by being proportional to ∣R∣|R|. Intuitively, ∣R∣|R| is a meaningful prior for the overall objectness of an input image. A diverse set of high-scoring non-overlapping regions are then picked from R′(c)R^{\prime}(c) as the pseudo boxes R^(c)\hat{R}(c) using non-maximum suppression. Even though being simple, this effective algorithm leads to significant performance improvements as shown in Sec. 5.

Bounding box regression is another module that plays an important role in supervised object detection but is missing in online self-training methods. To close the gap, we encapsulate a classification layer and a regression layer into ‘student blocks’ as shown via blue boxes in Fig. 2. We jointly optimize them using pseudo-labels R^\hat{R}. The predicted bounding boxes from the regression layer are referred to via μw(r)\mu_{w}(r) for all regions r∈Rr\in R. For each region rr, if it is highly overlapping with a pseudo-box r^∈R^\hat{r}\in\hat{R} for ground-truth class cc, we generate the regression target t^(r)\hat{t}(r) by using the coordinates of r^\hat{r} and by marking the classification label y^(c,r)=1\hat{y}(c,r)=1. The complete region-level loss for training the student block is:

where Lsmooth-L1\mathcal{L}_{\text{smooth-L1}} is the Smooth-L1 objective used in and λr\lambda_{r} is a scalar per-region weight used in .

In practice, conflicts happen when we force the y^(⋅,r)\hat{y}(\cdot,r) to be a one-hot vector since the same region can be chosen to be positive for different ground-truth classes, especially in the early stages of training. Our solution is to use that class for pseudo-label r^\hat{r} which has a higher predicted score s(c,r^)s(c,\hat{r}). In addition, the obtained pseudo-labels and the proposals are inevitably noisy. Imposing bounding box regression is able to correctly learn from the noisy labels by capturing the most consistent patterns among them, and refining the noisy proposal coordinates accordingly. We empirically verify in Sec. 5.3 that bounding box regression improves both robustness and generalization.

Self-ensembling.

We follow to stack multiple student blocks to improve performance. As shown in Fig. 2, the first pseudo-label R^1\hat{R}^{1} is generated from the teacher branch, and then the student block NN generates pseudo-label R^N\hat{R}^{N} for the next student block N+1N+1. This technique is similar to the self-ensembling method .

2 Concrete DropBlock

Because of the intra-category variation, existing WSOD methods often mistakenly only detect the discriminative parts of an object rather than its full extent. A natural solution for this issue encourages the network to focus on the context which can be achieved by dropping the most discriminative parts. Hence, spatial dropout is an intuitive fit.

Naïve spatial dropout has limition for detection since the discriminative parts of objects differ in location and size. A more structured DropBlock was proposed where spatial points on ROI feature maps are sampled randomly as blob centers, and the square regions around these centers of size H×HH\times H are then dropped across all channels on the ROI feature map. Finally, the feature values are re-scaled by a factor of the area of the whole ROI over the area of the un-dropped region so that no normalization has to be applied for inference when no regions are dropped.

By maximizing the original loss w.r.t. the Concrete DropBlock parameters, the Concrete DropBlock will learn to drop the most discriminative parts of the objects, as it is the easiest way to increase the training loss. This forces the object detector to also look at the context regions. We found this strategy to improve performance especially for non-rigid object categories, which usually have a large intra-class difference.

3 Sequential batch back-propagation

In this section, we discuss how we propose to handle memory limitations particularly during training, which turn out to be a major bottleneck preventing previous WSOD methods from using state-of-the-art deep nets. We introduce our memory-efficient sequential batch forward and backward computation, tailored for WSOD models.

Vanilla training via back-propagation stores all intermediate activations during the forward pass, which are reused when computing gradients of network parameters. This method is computationally efficient due to memoization, yet memory-demanding for the same reason. More efficient versions have been proposed, where only a subset of the intermediate activations are saved during a forward pass at key layers. The whole model is cut into smaller sub-networks at these key layers. When computing gradients for a sub-network, a forward pass is first applied to obtain the intermediate representations for this sub-network, starting from the stored activation at the input key layer of the sub-network. Combined with the gradients propagated from earlier sub-networks, the gradients of sub-network weights are computed and gradients are also propagated to outputs of earlier sub-networks.

This algorithm is designed for extremely deep networks where the memory cost is roughly evenly distributed along the layers. However, when these deep nets are adapted for detection, the activations (after ROI-Pooling) grow from 1×CHW1\times CHW (image feature) to N×CHWN\times CHW (ROI-features) where NN is in the thousands for weakly supervised models. Without ground-truth boxes, all these proposals need to be maintained for high recall and thus good performance (see the evidence in Appendix F).

To address this training challenge, we propose a sequential computation in the ‘Neck’ sub-module as depicted in Fig. 7. During the forward pass, the input image is first passed through the ‘Base’ and ‘Neck,’ with only the activation AbA_{b} after the ‘Base’ stored. The output of the ‘Neck’ then goes into the ‘Head’ for its first forward and backward pass to update the weights of the ‘Head’ and the gradients GnG_{n} as shown in Fig. 7 (a). To update the parameters of the ‘Neck,’ we split the ROI-features into ‘sub-batches’ and run back-propagation on each small sub-batch sequentially. Hence we avoid storing memory-consuming feature maps and their gradients within the ‘Neck.’ An example of this sequential method is shown in Fig. 7 (b), where we split 2000 proposals into two sub-batches of 1000 proposals each. The gradient GbG_{b} is accumulated and used to update the parameters of the ‘Base’ network via regular back-propagation as illustrated in Fig. 7 (c). For testing, the same strategy can be applied if either the number of ROIs or the size of the ‘Neck’ is too large.

Experiments

We assess our proposed method subsequently after detailing dataset, evaluation metrics and implementation.

We first conduct experiments on COCO , which is the most popular dataset used for supervised object detection but rarely studied in WSOD. We use the COCO 2014 train/val/test split and report standard COCO metrics including AP (averaged over IoU thresholds) and AP50 (IoU threshold at 50%).

We then evaluate on both VOC 2007 and 2012 , which are commonly used to assess WSOD performance. Average Precision (AP) with IoU threshold at 50% is used to evaluate the accuracy of object detection (Det.) on the testing data. We also evaluate correct localization accuracy (CorLoc.), which measures the percentage of training images of a class for which the most confident predicted box has at least 50% IoU with at least one ground-truth box.

Implementation details.

For a fair comparison, all settings of the VGG16 model are kept identical to except those mentioned below. We use 8 GPUs during training with one input image per device. SGD is used for optimization. The default pp and IoU in our proposed MIST technique (Alg. 1) are set to 0.150.15 and 0.20.2. For the Concrete DropBlock τ=0.3\tau=0.3. The ResNet models are identical to . Please check Appendix A for more details.

1 Overall performance

VGG16-COCO. We compare to state-of-the-art WSOD methods on COCO in Tab. 2. Our single model without any post-processing outperforms all previous approaches (w/ bells and whistles) by a great margin. On the private Test-dev benchmark, we increase AP50 by 11.2 (+82.3%). For the 2014 validation set, we increase AP and AP50 by 0.6 (+5.6%) and 1.6 (+7.1%). Complete results are provided in Appendix B. Note that compared to supervised models shown in the first two rows, the performance gap is still relatively big: ours is 56.9% of Faster R-CNN on average. In addition, our model achieves 12.4 AP and 25.8 AP50 on the COCO 2017 split as reported in Tab. 4, which is more commonly adopted in supervised papers.

ResNet-COCO. ResNet models have never been trained and evaluated before for WSOD. Nonetheless, they are the most popular backbone networks for supervised methods. Part of the reason is the larger memory consumption of ResNet. Without the training techniques introduced in Sec. 4.3, it’s impossible to train on a standard GPU using all proposals. In Tab. 2 we provide the first benchmark for the COCO dataset using ResNet-50 and ResNet-101. As expected we observe ResNet models to perform better than the VGG16 model. Moreover, we note that the difference between ResNet-50 and ResNet-101 is relatively small.

VGG16-VOC. To fairly compare with most previous WSOD works, we also evaluate our approach on the VOC datasets . The comparison to most recent works is reported in Tab. 3. All entries in this table are single model results. For object detection, our single-model results surpass all previous approaches on the publicly available 2007 test set (+1.3 AP50) and on the private 2012 test set (+1.9 AP50). In addition, our single model also performs better than all previous methods with bells and whistles (e.g., ‘+FRCNN’: supervised re-training, ‘+Ens.’: model ensemble). Combining the 2007 and 2012 training set, our model achieves 58.1% (+2.1 AP50) on the 2007 test set as reported in Tab. 4. CorLoc results on the training set and per-class results are provided in Appendix C. Since VOC is easier than COCO, the performance gap to supervised methods is smaller: ours is 78.1% of Faster R-CNN on average.

Additional training data. The biggest advantage of WSOD methods is the availability of more data. Therefore, we are interested in studying whether more training data improves results. We train our model on the VOC 2007 trainval (5011 images), 2012 trainval (11540 images), and the combination of both (16555 images) separately, and evaluate on the VOC 2007 test set. As shown in Tab. 4 (top), the performance increase consistently with the amount of training data. We verify this on COCO where 2014-train (82783 images) and 2017-train (128287 images) are used for training, and 2017-val (a.k.a. minival) for testing. Similar results are observed as shown in Tab. 4 (bottom).

2 Qualitative results

Qualitatively, we compare our full model with Tang et al. . In Fig. 9 we show a set of two pictures side by side, with baselines on the left and our results on the right. Our model is able to address instance ambiguity by: (1) detecting previously ignored instances (Fig. 9 left); (2) predicting tight and precise boxes for multiple instances instead of a big one (Fig. 9 center). Part domination is also alleviated since our model focuses on the full extent of objects (Fig. 9 right). Even though our model can greatly increase the score of larger boxes (see the horse example), the predictions may still be dominated by parts in some difficult cases.

More qualitative results are shown in Fig. 9 for all three datasets we used, as well in Appendix D. Our model is able to detect multiple instances of the same category (cow, sheep, bird, apple, person) and various objects of different classes (food, furniture, animal) in relatively complicated scenes. The COCO dataset is much harder than VOC as the number of objects and classes is bigger. Our model still tells apart objects decently well (Fig. 9 bottom row). We also show some failure cases (Fig. 9 right column) of our model which can be roughly categorized into three types: (1) relevant parts are predicted as instances of objects (hands and legs, bike wheels); (2) in extreme examples, part domination remains (model converges to a face detector); (3) object co-occurrence confuses the detector when it predicts the sea as a surfboard or the baseball court as a bat.

3 Analysis

How much does each module help? We study the effectiveness of each module in Tab. 6. We first reproduce the method of Tang et al. , achieving similar results (first two rows). Applying the developed MIST module improves the results significantly. This aligns with our observation that instance ambiguity is the biggest bottleneck for WSOD. Our conceptually simple solution also outperforms an improved version (PCL), which is based on a computationally expensive and carefully-tuned clustering.

The devised Concrete DropBlock further improves the performance when using MIST as the basis. This module surpasses several variants including: (1) (Img Spa.-Dropout): spatial dropout applied on the image-level features; (2) (ROI-Spa.-Dropout): spatial dropout applied on each ROI where each feature point is treated independently. This setting is similar to ; (3) (DropBlock): the best-performing DropBlock setting reported in .

To validate that instance ambiguity is alleviated, we report Average Recall (AR) over multiple IoU values (.50:.05:.95.50:.05:.95), given 1, 10, 100 detections per image (AR1AR^{1}, AR10AR^{10}, AR100AR^{100}) and for small, medium, annd large objects (ARsAR^{\text{s}}, ARmAR^{\text{m}}, ARlAR^{\text{l}}) on VOC 2007. We compare the model with and without MIST in Tab. 6 where our method increases all recall metrics.

Has Part Domination been addressed?

In Fig. 11, we show the 5 categories with the biggest relative performance improvements on the VOC 2007 and VOC 2012 dataset after applying the Concrete DropBlock. The performance of animal classes including ‘person’ increases most, which matches our intuition mentioned in Sec. 1: the part domination issue is most prominent for articulated classes with rigid and discriminative parts. Across both datasets, three out of the five top classes are mammals.

Space-time analysis of sequential batch BP?

We also study the effect of our sequential batch back-propagation. We fix the input image to be of size 600×600600\times 600, and run two methods (vanilla back-propagation and ours with sub-batch size 500 using ResNet-101 for comparison. We change the number of proposals from 1k to 5k in 1k increments, and report average training iteration time and memory consumption in Fig. 11. We observe: (1) vanilla back-propagation cannot even afford 2k proposals (average number of ROIs widely used in ) on a standard 16GB GPU, but ours can easily handle up to 4k boxes; (2) the training process is not greatly slowed down, ours takes ∼\sim1-2×\times more time than the vanilla version. In practice, input resolution and total number of proposals can be bigger.

Robustness of MIST?

To assess robustness we test a baseline model plus this algorithm only using different top-percentage pp and rejection IoU on the VOC 2007 dataset. Results are shown in Fig. 12. The best result is achieved with p=0.15p=0.15 and IoU=0.2IoU=0.2, which we use for all the other models and datasets. Importantly, we note that, overall, the sensitivity of the final results on the value of pp is small and only slightly larger for IoU.

4 Extension: video object detection

We finally generalize our models to video-WSOD, which hasn’t been explored in the literature. Following supervised methods, we experiment on the most popular dataset: ImageNet VID . Frame-level category labels are available during training. Uniformly sampled key-frames are used for training following and evaluation settings are also kept identical. Results are reported in Tab. 7. The performance improvement of the proposed MIST and Concrete DropBlock generalize to videos. The memory-efficient sequential batch back-propagation permits to leverage short-term motion patterns (i.e., we use optical-flow following ) to further increase the performance. This suggests that videos are a useful domain where we can obtain more data to improve WSOD. Full details are provided in Appendix G.

Conclusion

In this paper, we address three major issues of WSOD. For each we have proposed a solution and demonstrated its effectiveness through extensive experiments. We achieve state-of-the-art results on popular datasets (COCO, VOC 07 and 12) and are the first to benchmark ResNet backbones and weakly supervised video object detection.

Acknowledgement: ZR is supported by Yunni & Maxine Pao Memorial Fellowship. This work is supported in part by NSF under Grant No. 1718221 and No. 1751206.

References

Appendix A Implementation Details

In this section, we provide additional implementation details for completeness.

We use the standard VGG-16 (without batch normalization) as backbone. As shown in Fig. 2, the ‘Base’ network contains all the convolutional layers before the fully-connected layers. Following , we remove the last max-pooling layer, and replace the penultimate max-pooling layer and the subsequent convolutional layers with dilated convolutional (dilation=2) layers to increase the feature map resolution. Standard RoI-pooling is used for computing region-level features. We use the fully-connected layers of VGG-16 except the last classifier layer as the ‘Neck’. After ‘Neck’, the RoI features are projected to fw,gw,s^w,μ^wf_{w},g_{w},\hat{s}_{w},\hat{\mu}_{w} using 4 single fully-connected layers.

ResNets

We use the ResNet-50/101-C4 variant from Detectron code repository . Convolutional layers of the first 4 ResNet stages (C1-C4) are used as ‘Base’ and the last stage (C5) is used as ‘Neck’. Standard RoI-pooling is used, and RoI features are projected using linear layers.

A.2 Concrete DropBlock

Concrete DropBlock is implemented as a standard residual block as in ResNets. It takes as input the RoI features and output a 1 channel heatmap pθ(r)p_{\theta}(r). On the skip connection we use 1×11\times 1 convolution to reduce feature channels. We then generate the hard mask Mθ(r)M_{\theta}(r) using Gumbel-softmax, and the structured dropout region as in DropBlock .

A.3 Student Blocks

Following , we stack 3 student blocks. During training, student block NN generates pseudo labels for the next student block N+1N+1. During testing, we average the predictions of all student blocks as final results.

A.4 Training

Our code is implemented in PyTorch and all the experiments are conducted on single 8-GPU (NVIDIA V100) machine. SGD is used for optimization with weight decay 0.0001 and momentum 0.9. The batch size and initial learning rate is set to 8 and 0.01 on VOC 2007; 16 and 0.02 on VOC 2012. On both datasets we train the model for 30k iterations and decay the learning rate by 0.1 at 20k and 26k steps. On COCO, we train the model for total 130k iterations and decay the learning rate at 90k and 120k steps with batch size 8 and initial learning rate 0.01. We use Selective-Search (SS) for VOC datasets and MCG for COCO.

A.5 Data Augmentation & Inference

Multi-scale inputs (480, 576, 688, 864, 1000, 1200) are used during both training and testing following and the longest image side to set to less than 2000. At test time, the scores are averaged over all scales and their horizontal flips.

Appendix B Additional quantitative results on COCO

In Tab. 8, we report quantitative results at different thresholds and scales on COCO for different models. The reported metrics include: Average Prevision (APAP) over multiple IoU thresholds (.50 : .05 : .95), at IoU threshold 50% and 75% (AP50AP^{50}, AP75AP^{75}), and for small, medium and large objects (APsAP^{\text{s}}, APmAP^{\text{m}}, APlAP^{\text{l}}); and Average Recall (ARAR) over multiple IoU values (.50 : .05 : .95), given 1, 10 and 100 detections per image (AR1,AR10,AR100AR^{1},AR^{10},AR^{100}); and for small, medium and large objects (ARsAR^{\text{s}}, ARmAR^{\text{m}}, ARlAR^{\text{l}}). The results in Tab. 8 show that object size is a significant factor that influences the detection accuracy. The detector tends to perform better on large objects rather than smaller ones.

Appendix C Additional results on VOC

In Tab. 10 and Tab. 10, we report the per-class detection APs on the test sets of both VOC 2007 and 2012. Compared to other WSOD methods we observe: (1) Our method outperforms all others on most categories (10 classes on VOC 2007, 14 classes on VOC 2012). (2) The classes that are hard for our approach (e.g., boat, plant, and chair) are also challenging for other methods. This suggests that these categories are essentially hard examples for WSOD methods, for which a certain amount of strong supervision might still be needed.

Compared to supervised models (Fast R-CNN, Faster R-CNN) we note: (1) Our weakly supervised model performs competitively for classes such as: airplane, bicycle, bus, car, cow, motorbike, sheep, tv-monitor, where the performance gap is usually less than 10% AP. Our model sometimes even outperforms supervised models on categories that are considered relatively easy with small intra-class difference (bicycle and motorbike in VOC 2007, motorbike and tv-monitor in VOC 2012). (2) For classes like boat, chair, dinning table, person, all WSOD methods are significantly worse than supervised methods. This is likely due to a large intra-class variation. WSOD methods fail to capture the consistent patterns of these classes.

C.2 Per-class correct localization results

In Tab. 11 and Tab. 12, we report the per-class correct localization (CorLoc) results on the trainval sets of both VOC 2007 and VOC 2012. Consistent with prior work this metric is computed on the training set. Thus it does not reflect the true performance of the detection models and has not been widely adopted by supervised methods . For WSOD approaches, it serves as an indicator of the ‘over-fitting’ behavior. Compared with previous state-of-the-art, our method achieves the third best result on VOC 2007, winning on 2 categories. We also achieve the second best performance on VOC 2012 and win on 19 categories. We find that: (1) Our model performs well for classes like: airplane, bicycle, bottle, bus, motorbike, sheep, tv-monitor. This observation aligns very well with the detection results. (2) The best performing methods differ across classes, which suggest that methods could potentially be ensembled for further improvements.

Appendix D Additional qualitative results

We show additional results that highlight cases of ‘Instance Ambiguity’ and ‘Part Domination’ in Fig. 13 and Fig. 14, respectively. Following the main paper, we compare our final model to a baseline without the modules proposed in Sec. 4.1 and Sec. 4.2 of the main paper to demonstrate the effectiveness of these two modules visually. We show a set of two pictures side by side, the baseline on the left and ours on the right. From the results, we observe: (1) we have addressed the ‘Missing Instances’ issue and previously ignored objects are detected with great recall (e.g., monitor, sheep, car, and person in Fig. 13); (2) we have addressed the ‘Grouped Instances’ issue as our model predicts tight and precise boxes for multiple instances rather than one big one (e.g., bus, motor, boat, car in Fig. 13); (3) we have also alleviated the ‘Part Domination’ issue for objects like dog, cat, sheep, person, horse, and sofa (see Fig. 14).

We also provide additional visualization of our results on COCO in Fig. 15. We obtain these results by running the VGG16 based model on the COCO 2014 validation set. Our model is able to detect different instances of the same category (e.g., car, elephant, pizza, cow, umbrella) and various objects of different classes in relatively complicated scenes, and the obtained boxes can cover the whole objects pretty well rather than simply focusing on discriminative parts.

D.2 Results on ImageNet VID dataset

Additional visualizations of our obtained results on ImageNet VID are shown in Fig. 16, where the frames of the same video are illustrated in the same row. These results are obtained using the ResNet-101 based model. We observe: our model is able to handle objects of different poses, scales, and viewpoints in the videos.

Appendix E Proposal statistics

For consistency with prior literature, we use Selective-Search (SS) for VOC and MCG for COCO. Both methods generate around 2K proposals on average as shown in Tab. 13 but occasionally yield more than 5K on certain images. Our Sequential batch back-propagation can handle these cases easily even with ResNet-101, while other methods quickly run out of memory (Fig. 11 in main paper).

Appendix F Need for redundant proposals

In WSOD, since ground-truth boxes are missing, object proposals have to be redundant for high recall rates, consuming significant amounts of memory. To study the need for a large number of proposals we randomly sample pp percent of all proposals. A VGG16 based model on VOC 2007 is used. The results are summarized in Tab. 14. Reducing the number of proposals even by a small amount significantly reduces accuracy: using 95% of the proposals causes a 2.8% AP drop. This suggests that all proposals should be used for best performance.

Appendix G Additional details on video experiments

In this section, we provide additional details of Sec. 5.4. Following supervised methods for video object detection , we experiment on the most popular dataset: ImageNet VID . Frame-level category labels are available during training. For each video, we use the uniformly sampled 15 key-frames from for training. For evaluation, we test on the standard validation set, where per-frame spatial object detection results are evaluated for all the videos.

The two models ‘Ours’ and ‘Ours (MIST only)’ are two single-frame baselines with or without Concrete DropBlock (main paper Sec. 4.2). In addition, the memory-efficient sequential batch back-propagation (main paper Sec. 4.3) permits to leverage short-term motion patterns (i.e., optical-flow) to further increase the performance. For ‘Ours+flow,’ we first use FlowNet2 to compute optical flow between neighboring frames and the reference frame. The estimated flow maps are then used to warp the nearby frames’ feature maps to linearly sum with the reference frame for representation enhancement. The accumulated features are then fed into the proposed task head (modules after ‘Base’ in main paper Fig. 2) for weakly supervised training. This method combines the flow-guided feature warping method as discussed in to leverage temporal coherence and the proposed WSOD task head to handle frame-level weak supervision. Hence it achieves better results than the aforementioned two baselines (‘Ours’ and ‘Ours (MIST only)’) using both VGG16 and ResNet-101 as reported in Tab. 7.