Boosting Weakly Supervised Object Detection via Learning Bounding Box Adjusters

Bowen Dong, Zitong Huang, Yuelin Guo, Qilong Wang, Zhenxing Niu, Wangmeng Zuo

Introduction

Object detection has attracted considerable attention in computer vision community, and benefits a wide range of applications. Along with the development of powerful convolutional neural networks (CNNs) and large-scale well-annotated datasets, the performance of object detection networks has achieved remarkable improvement. Nevertheless, the success of object detection networks highly depends on precise but costly instance-level bounding box annotations of abundant images. To alleviate this issue, weakly supervised object detection (WSOD) aiming at learning effective detection models with image-level supervision has emerged as an inspiring recent topic.

Existing WSOD methods usually adopt the multiple instance learning (MIL) framework based on the precomputed proposals. And most efforts have been given to improve proposal classification ability. However, the bounding boxes of most existing methods are mainly determined by precomputed proposals, thereby being limited in precise object localization. For single-phase WSOD methods , the precomputed proposals classified to a specific class are directly taken as the detection results. Bounding box regression branches are introduced in and multi-phase training are adopted in . But they are usually supervised based on the pseudo ground-truths by selecting precomputed proposals with the highest scores. In terms of localization performance, there remains a huge gap between WSOD methods and their fully-supervised counterparts.

Transfer learning has also been investigated to improve the localization performance of WSOD. Lee et al. presented a universal bounding box regressor (UBBR) trained on a well-annotated auxiliary dataset for refining bounding boxes generated in WSOD. Instead, Uijlings et al. trained a universal detector on the well-annotated source dataset, which is then transferred to WSOD as a generic proposal generator. However, and adopt the single-stage transfer strategy, which actually are not specified to WSOD and suffer from imperfect annotations in source domain . Going beyond , Zhong et al. trained and exploited the one-class universal detector (OCUD) in a progressive manner. In contrast, both the source well-annotated and target weakly annotated datasets are required in the whole training process for OCUD . When the source dataset is private and is of large scale , it is preferred to avoid the direct joint use of the source and target datasets for WSOD with transfer learning. Instead, the owner of source datasets can first extract knowledge from data and then distribute knowledge instead of source datasets to the user for boosting WSOD.

In this paper, we follow the problem setting in , and propose a learnable bounding box adjuster (LBBA) for boosting WSOD performance. Specifically, we consider a well-annotated auxiliary dataset and a weakly annotated dataset. Our method involves two subtasks, i.e., learning class-agnostic bounding box adjuster and training LBBA-boosted WSOD model. In comparison to , the LBBAs are specifically designed for improving WSOD performance by developing a multi-stage scheme. Different from , only the LBBAs and weakly-annotated dataset are used for boosting WSOD, and thus our approach is practically convenient and economical for WSOD training while avoiding the leakage of the auxiliary dataset.

To better learn LBBAs from the well-annotated auxiliary dataset and exploit them to improve the performance of WSOD, we formulate the learning of LBBAs as a bi-level optimization problem and present an EM-like multi-stage training algorithm. In particular, the lower subproblem is formulated to learn a deep detection model by incorporating WSOD with LBBA-based regularization, while the upper subproblem is formulated to learn the boundary box adjuster for regressing the selected region proposals generated by WSOD towards the ground-truth bounding boxes. With such formulation, the LBBAs can thus be learned for optimizing WSOD performance. For solving the bi-level optimization problem, we adopt an EM-like multi-stage training algorithm by alternating between training LBBA and WSOD models. Given the class-agnostic and multi-stage LBBAs, the training of LBBA-boosted WSOD also involves several stages. In each stage, the final LBBA can be used to predict the bounding boxes based on the selected region proposals generated by WSOD, which are then used to train the WSOD models.

Nevertheless, our LBBAs improve localization performance but are limited in improving proposal classification. As a remedy, we introduce a masking strategy to improve the classification performance of the detector. Specifically, a multi-label classifier is introduced to predict category confidence on image-level, which can further suppress scores of false-positive proposals of WSOD network.

Extensive experiments have been conducted to evaluate our proposed method. Benefiting from the class-agnostic setting, LBBAs generalize well to new classes of objects and improves the localization performance of WSOD. Our method performs favorably against state-of-the-art WSOD methods as well as knowledge transfer models with similar problem setting, e.g., UBBR . Contributions of this work can be summarized as follows:

Multi-stage learnable bounding box adjusters are presented for improving localization performance of WSOD, which is the core component of our proposed framework. Particularly, LBBAs make it feasible to use source and target datasets separately for training WSOD models, which is practically more convenient and economical.

A bi-level optimization formulation, as well as an EM-like multi-stage training algorithm, are suggested to learn LBBAs specified for optimizing WSOD.

An effective masking strategy is introduced to improve the accuracy of the proposal classification branch.

Experimental results show our proposed method performs favorably against the state-of-the-art WSOD methods and knowledge transfer models with the similar problem setting.

Related Work

Weakly supervised object detection (WSOD) aims at training an effective detector only using image-level labels, and is usually formulated as a multiple instance learning (MIL) problem . Existing WSOD approaches can be roughly grouped into two categories: single-phase training methods and multi-phase training ones. For single-phase training methods, they rely on precomputed proposals during training and testing. Specifically, Bilen et al. proposed a two-stream detection network (WSDDN) as the basic proposal classifier. To improve proposal classification ability, OICR and PCL proposed online classifier refinement module. OIM proposed spatial and appearance graphs with object instance reweighted loss to resolve part domination. SDCN and WS-JDS introduced segmentation branch and collaboration loop to reweight proposals. As for improving proposal localization ability, Yang et al. , WSOD2 and MIST introduced bounding box regression into WSOD network, where proposals with highest scores are selected as pseudo ground-truths to supervise bounding box regression branch.

For multi-phase training methods , an additional detector is further trained by selecting proposals with the highest scores as pseudo ground-truths based on the output of trained WSOD network in the prior phase . Any single-phase methods can be extended to multi-phase setting by this procedure. Current multi-phase training methods focus on how to select pseudo ground-truths with the highest scores. However, these approaches rely on only selected precomputed proposals to localize objects or supervise box regression branch, low precision proposals restrict the localization ability of WSOD approaches. Different from the above methods, we aim at resolving this issue by using learnable bounding box adjusters, which provide more precise pseudo boxes supervision to help WSOD network obtain better object localization ability.

2 Transfer Learning in WSOD

Transfer learning based WSOD usually leverages an auxiliary dataset to provide semantic information or class-agnostic information to help WSOD networks train on weakly-annotated target dataset. Previous works focused on transferring semantic information between strong classifier and weakly supervised detector. Among them, Hoffman et al. proposed LSDA, which introduces category specific adaptation to adapt a classifier into target detection dataset. Tang et al. further extended LSDA by building visual similarity and semantic relatedness. Nonetheless, above methods are not proposed for improving bounding box regression.

Recently, several approaches have been studied to exploit transfer learning for improving object localization performance. proposed to learn proposal generators to help WSOD network locate novel objects on weakly-annotated target dataset. Among them, trained proposal generators merely using the auxiliary dataset, while Zhong et al. trained generator on both auxiliary dataset and weakly-annotated dataset progressively to generalize better on target dataset. Instead, Lee et al. proposed a box refinement module, which takes the random transformations of ground-truth boxes as the input to learn class-agnostic box regressor, and also exhibits certain generalization ability on target weakly-annotated dataset. However, the real boxes generated during WSOD training may be quite different from those by random transformations, making the learned regressor not tailored to WSOD. In comparison to existing methods, our LBBAs can be considered as the multi-stage training of box refinement modules only using the auxiliary dataset, and achieves very competitive box regression performance on weakly-annotated dataset. Different from UBBR, our method dynamically takes the proposals generated by WSOD as the input to train LBBA, and thus is expected to achieve improved detection performance.

Proposed Method

2 Overview

3 Baseline WSOD Model

where BCE(⋅,⋅)\text{BCE}(\cdot,\cdot) denotes the binary cross-entropy loss. To improve detection quality, we also introduce pseudo label mining strategy and construct instance refinement branch optimized by a set of weighted instance refinement loss Lr\mathcal{L}_{\text{r}} .

where Lr\mathcal{L}_{\text{r}} and Lrpn-cls\mathcal{L}_{\text{rpn-cls}} are the cross-entropy losses supervised by pseudo class labels on the selected proposals, while Lrpn-det\mathcal{L}_{\text{rpn-det}} and Ldet\mathcal{L}_{\text{det}} are the smooth-L1 losses supervised by the proposal boxes of pseudo ground-truths. Note that we follow the same strategy of OICR to generate pseudo ground-truths.

We note that the bounding box regression branch in baseline WSOD model is learned based on the supervision from the precomputed proposals, which naturally are not precise enough. In the subsequent subsections, we learn a set of bounding box adjusters to provide better ground-truth for supervising the bounding box regression branch, thereby being beneficial to detection performance. Moreover, we use the above baseline WSOD model as an example to show the effectiveness of the learned bounding box adjusters. Actually, our proposed method is independent with most existing WSOD methods and can be incorporated with them to further boost detection performance. And we will illustrate this point in the experiments.

4 Learning Bounding Box Adjusters

To formulate our weakly supervised object detection problem elegantly, we first revisit the traditional EM algorithm for weakly supervised learning. In particular, E-step is used to update latent variable b^\hat{\text{b}},

where L\mathcal{L} is a combination of weakly supervised object detection loss Lwsod\mathcal{L}_{\text{wsod}} and bounding box regression loss Lbbr\mathcal{L}_{\text{bbr}}.

After introducing LBBA gg into WSOD, our WSOD problem can be transferred into a bi-level optimization problem, here we state how to build bi-level optimization.

where gg generates adjusted bounding box regression for given proposals from WSOD fauxf^{\text{aux}}. Thus upper subproblem has transferred into a fully-supervised setting.

4.2 EM-like Multi-stage Training Algorithm

To sum up, after the initialization, our training algorithm alternates between the E-step and M-step for TT times. Hence, it is a multi-stage training scheme, where we run the E-step and M-step once in each stage. The training process of LBBA is given in Algorithm 1.

5 LBBA-boosted WSOD

Nonetheless, we empirically find that updating WSOD network with only the last gTg_{T} can attain a similar performance. Hence we can build a lighter pipeline by only using the last gTg_{T}.

6 Masking Strategy for Proposal Classification

Experiments

Auxiliary Dataset. MS-COCO 2017 is a large-scale object detection dataset. Note that MS-COCO dataset includes 80 different object classes. To eliminate semantic overlap and show the generalization ability of our method, we construct a subset of MS-COCO by excluding PASCAL VOC classes instance annotations and call it COCO-60. As such, COCO-60 dataset contains ∼\sim98K training images and ∼\sim4K validation images, respectively.

Target Datasets. PASCAL VOC 2007 and 2012 datasets contain 9,963 images and 22,531 images collected from 20 object classes. For fair comparison, we use trainval set for training WSOD networks and report evaluation results on test set. During the training process, only image-level labels are used as supervision. We also utilized other datasets to evaluate our LBBA, see the suppl. for details.

Evaluation Metrics. Since our method aims at improving object detection performance, Average Precision (AP) is used as the basic evaluation metric in our experiments. We also adopt CorLoc as another evaluation metric.

2 Comparison with State-of-the-arts

We state the implementation details in the suppl. and we build up all experiments based on it. We compare our method with several state-of-the-art WSOD approaches in terms of detection and localization performance on PASCAL VOC datasets. As suggested in , we report detection results on test set and localization results on trainval set, respectively. Table 1 compares the results of different state-of-the-art WSOD approaches on PASCAL VOC 2007 and 2012 datasets. It can be seen that our LBBA improves OICR and OICR+REG over 15.3% and 5.0% on PASCAL VOC 2007 dataset, respectively. Furthermore, our method performs better than all competing methods, except Zhong et al. . Note that uses stronger backbone model and knowledge transfer strategy by directly incorporating source and target datasets. Moreover, the auxiliary dataset adopted in Zhong et al. is different from ours (See the suppl. for more details). As shown in Fig. 2, our method has the ability to generate precise bounding boxes. On PASCAL VOC 2012, our LBBA is superior to all competing methods and obtains more than 1% gains over all WSOD approaches. Experimental results show that our method is effective in improving the detection performance of WSOD.

We further evaluate the localization performance of our method. Table 2 lists results of several state-of-the-art WSOD approaches on PASCAL VOC 2007 and 2012. Our LBBA outperforms OICR by 11.7% and also improves the baseline OICR+REG over 4.3% on PASCAL VOC 2007 dataset. Besides, our LBBA performs better than all competing methods. Meanwhile, on PASCAL VOC 2012, our LBBA is also superior to all competing methods, and obtains 1.3% gain over WSOD 2. In comparison to Zhong et al. , our LBBA-based method employs a weaker backbone model and avoids the direct joint use of the source and target datasets, while still achieving competitive CorLoc results under the settings of both single-scale testing and multi-scale testing. The above results show that our LBBA-based method is effective in improving the localization performance of WSOD.

3 Ablation Study

Additionally, we employ PASCAL VOC 2007 to assess the effect of some key components on our LBBA. We state a more detailed ablation study in the suppl..

Backbone Models of Adjuster gg. In this work, Faster R-CNN is used as adjuster. Here, we first evaluate the effect of backbone models on adjuster gg. To this end, we compare two CNN architectures as backbone models of Faster R-CNN, i.e., ResNet-50 and VGG-16 . Particularly, we set iterations TT of multi-stage learning to 33 and adopt WSDDN as WSOD network ff. The compared results on VOC 07 are listed in Table 3, from which we can see that adjuster gg with backbone of ResNet-50 outperforms one with backbone of VGG-16 by 2.5% and 2.6% in terms of mAP and CorLoc, respectively. These results show that our method can benefit from a stronger adjuster, which encourages us to develop more effective adjusters.

Effect of WSOD network ff. After determining backbone model of adjuster gg, we access the impact of WSOD network ff. Specifically, we consider three methods (i.e., WSDDN+REG , OICR+REG and OICR+REG with top p% pseudo label mining ) for our WSOD network ff, and compare our LBBA with the original methods (i.e., baseline). The iterations TT of multi-stage learning is set to 33, and the results of different WSOD networks ff are given in Table 4. First, our LBBA achieves clear performance gains (more than 4%) over the baseline methods for all choices of WSOD networks in terms of mAP and CorLoc. It demonstrates that the proposed LBBA methods can be well generalized to various WSOD networks. Second, our LBBA benefits from stronger WSOD networks, and so we compare with state-of-the-arts by using OICR+ as WSOD network ff.

Multi-stage LBBAs. The proposed multi-stage learning strategy of LBBAs involves two core factors, i.e., number of iterations (TT) and learnable, auxiliary WSOD network fauxf^{\text{aux}}. By fixing WSOD network ff and adjuster gg respectively be OICR+ and Faster R-CNN with backbone of ResNet-50, we assess the effects of number of iterations (TT) and fauxf^{\text{aux}} on our LBBA method. To this end, we learn bounding box adjusters by setting TT from 0 to 3. Besides, we replace learnable fauxf^{\text{aux}} by using MCG to generate proposals, namely LBBA-MCG. Table 5 gives the results of adjuster gg and WSOD network ff on COCO-60 and VOC 07 using different learning strategies, respectively. It can be seen that increasing iterations (TT) can improve performance of both adjuster gg and WSOD network ff. However, performance of WSOD network ff is sightly improved, when number of iterations T>2T>2. Therefore, T=3T=3 is a good choice to balance efficiency and effectiveness. These results clearly demonstrate the effectiveness of our multi-stage learning strategy. The learnable fauxf^{\text{aux}} with 3 iterations is superior to LBBA-MCG by 1.3% and 1.5% for adjuster gg and WSOD network ff, showing the significance of learnable fauxf^{\text{aux}}.

Conclusion

In this paper, we presented a knowledge transfer based WSOD method. Our proposed method involves two subtasks, i.e., learning bounding box adjusters and LBBA-boosted WSOD. For the former subtask, we suggested a bi-level optimization formulation on the auxiliary dataset and an EM-like training algorithm to learn multi-stage and class-agnostic LBBAs specified for optimizing WSOD performance. For the later subtask, we adopted a multi-stage scheme to utilize only the LBBAs and weakly-annotated dataset for WSOD. Additionally, a masking strategy is adopted to improve proposal classification for benefiting detection performance. Experimental results show that our proposed method performs favorably against the state-of-the-art WSOD methods and knowledge transfer model with similar problem setting . Nonetheless, we mainly focus on transferring across classes in this paper, while the transferring across domains is not specifically considered. In the future, we will explore suitable domain generalization methods for coping with this issue.

Acknowledgement

This work was supported in part by the National Natural Science Foundation of China under grant No.s U19A2073 and 61806140, and Natural Science Foundation of Tianjin under grant No. 20JCQNJC1530.

References

Appendix A Discussion of EM-like training algorithm

The reason why EM-like training is necessary is that the problem is formulated as a bi-level optimization problem, direct joint training to solve this problem is harmful to the generalization ability of LBBA. And EM-like training can keep that of LBBA. Here we state why formulating WSOD problem as bi-level optimization.

In particular, E-step is used to update latent variable b^\hat{\text{b}},

where L\mathcal{L} is a combination of weakly supervised object detection loss Lwsod\mathcal{L}_{\text{wsod}} and bounding box regression loss Lbbr\mathcal{L}_{\text{bbr}}.

After introducing LBBA gg into WSOD, our WSOD problem can be transferred into a bi-level optimization problem, here we state how to build bi-level optimization.

where gg generates adjusted bounding box regression for given proposals from WSOD fauxf^{\text{aux}}. Thus upper subproblem has transferred into a fully-supervised setting. Furthermore, to ease the training difficulty of the upper subproblem and improve the precision of b^aux\hat{\text{b}}^{\text{aux}}, we modify the upper subproblem by requiring LBBA accurately predicts the ground-truth boxes, resulting in the following bi-level optimization formulation.

Appendix B Datasets

To illustrate the effectiveness of our method, we conduct experiments on various representative datasets: PASCAL VOC 2007 and 2012 datasets, MS-COCO dataset, and ILSVRC 2013 detection dataset.

MS-COCO 2017 is a large-scale object detection dataset. Note that MS-COCO dataset includes 80 different object classes. To eliminate semantic overlap and show the generalization ability of our method, we construct a subset of MS-COCO by excluding PASCAL VOC classes instance annotations and call it COCO-60. As such, COCO-60 dataset contains ∼\sim98K training images and ∼\sim4K validation images, respectively. Construction details are shown as Appendix B.4.

To prove that our method can be generalized to more categories, we conduct extended experiments on the ILSVRC2013 detection dataset. ILSVRC detection dataset contains 200 categories, which is much more than that for PASCAL VOC or COCO-20. To construct the corresponding auxiliary dataset, we select the first 100 classes sorted in alphabetic order as the source classes in the auxiliary dataset. Construction details are shown as Appendix B.5.

B.2 Target Datasets

PASCAL VOC 2007 and 2012 datasets contain 9,963 images and 22,531 images collected from 20 object classes, respectively. For fair comparison, we use trainval set for training WSOD networks and report evaluation results on test set. During the training process, only image-level labels are used as supervision.

To verify the generalization ability of our LBBA, we construct another target dataset from MS-COCO dataset namely COCO-20 dataset. Note that the COCO-20 dataset has the same 20 classes as PASCAL VOC dataset, but containing more complicated scenarios in images. Construction details are shown as Appendix B.5.

ILSVRC detection dataset contains 200 categories. To construct the target dataset and avoid semantic overlaps with the corresponding auxiliary dataset, we select the last 100 classes sorted in alphabetic order as target classes in our weakly supervised object detection dataset. Construction details are shown as Appendix B.5.

B.3 Auxiliary-Target Pairs

From these datasets, we divide them into four dataset-pair settings, an auxiliary dataset corresponding to a target dataset, to deploy experiments. Table 11 give the dataset-pair settings. Setting 1 and Setting 2 are mentioned in section 4 of main paper and we will introduce details of setting 3 and setting 4 in Appendix B.4 and Appendix B.5. Then we will state more experimental results in Appendix F and Appendix G.

B.4 Construction of COCO-60/COCO-20

To simplify the statement, we define COCO-60 classes as the categories in original COCO classes but excluding PASCAL VOC classes. Then we state how to construct COCO-60 dataset and COCO-20 dataset.

To construct COCO-60 dataset, we first keep annotations of COCO-60 classes in COCO 2017 train set, then we select images which contain at least one instance of COCO-60 classes in COCO 2017 train set to construct our COCO-60 train set. Next we keep the same steps to build up our COCO-60 val set.

Besides, we also follow Zhong et al. to define a COCO-60-clean dataset. Particularly, we select images which only contain instances of COCO-60 classes in COCO 2017 train set to construct COCO-60-clean train set, and obtain only 21987 training images. Compared to COCO-60 dataset, COCO-60-clean dataset does not exist objects of VOC classes in the background of images, such that this dataset is cleaner than our COCO-60 dataset and easier to learn. We will discuss the difference between our method and Zhong et al. based on COCO-60 and COCO-60-clean datasets.

As for COCO-20 dataset, we select images which only contain instances of 20 PASCAL VOC classes in COCO 2017 train set to construct our COCO-20 train set. Next we keep annotations of 20 PASCAL VOC classes in COCO 2017 val set, and then select images which contain at least one instance of 20 PASCAL VOC classes in COCO 2017 val set to construct our COCO-20 val set.

B.5 Construction of ILSVRC-Source/Target

The original ILSVRC dataset contains a training set and a validation set. Firstly, We split the validation set into val1 validation set and val2 validation set. Then we state how to construct ILSVRC-Source dataset and ILSVRC-Target dataset.

To construct ILSVRC-Source training set, we keep images of the first 100 categories sorted in alphabetic order from val1 and sample 1000 images per category in the same 100 categories from ILSVRC training set as data augmentation.

To construct ILSVRC-Target training set, we keep images of the latter 100 categories sorted in alphabetic order from val1 and sample a maximum of 1000 images per category in latter categories from ILSVRC training set to augment it,while keeping only image-level labels. And to construct ILSVRC-Target test set, we keep images of the same 100 categories from val2.

Appendix C Implementation Details

For LBBAs, we apply Faster R-CNN with backbone of ResNet-50 and we adopt class-agnostic bounding box adjusters to eliminate potential semantic information leak in bounding box refinement. For WSOD network, we apply OICR with a backbone of VGG-16 and introduce a class-agnostic bounding box regression branch. Following the settings of , we initialize backbone models of two networks with ImageNet pre-trained weights while other layers are randomly initialized. As suggested in , we use MCG boxes as precomputed proposals for COCO-60 and use Selective Search boxes as precomputed proposals for PASCAL VOC. During training, both two networks are optimized by stochastic gradient descent (SGD) with the batch size of 1 and initialized learning rate of 0.001. In each stage, LBBA is trained with 4 epochs, and the learning rate is decayed by 0.1 after 3 epochs. Analogously, WSOD network is trained within 20 epochs and learning rate is decayed by 0.1 after 10 epochs. All programs are implemented by PyTorch toolkit, and all experiments are conducted on a single NVIDIA RTX 2080Ti GPU.

For the multi-label image classifier, we adopt the ADD-GCN , which builds a Dynamic Graph Convolutional Network (D-GCN) to model the relation of content-aware category representations generated by a Semantic Attention Module(SAM). During training, the ADD-GCN is optimized by SGD with batch size of 16. The learning rate is initially set to 0.05 for training 40 epoch and decayed by 0.1 to train the latter 10 epoch. The best threshold τ\tau is set to -3.0. By the way, the setting of the τ\tau is based on the implementation of multi-label image classifier. Too high or too low will be detrimental to the final result, and we will give the results and analysis in the next section.

All the source code and pre-trained models will be made publicly available.

C.2 Structure of LBBA

Here we briefly introduce the structure of LBBA. In our solution, we adopt Faster R-CNN with backbone of ResNet-50 as our LBBA. And LBBA is designed to be a class-agnostic bounding box regressor to eliminate potential semantic information leak in bounding box refinement. Note that the inside RPN is only used during EM-like LBBA training to improve the training stabilization and generalization ability of LBBA, and will not be used during the inference stage. We argue that using Faster R-CNN as adjuster has two merits. (i) For the initialization of LBBA training, Faster R-CNN exhibits better performance than Fast R-CNN. (ii) By combining precomputed proposals and proposals from RPN, box regression branch of LBBA can generalize better to various proposals, resulting in more precise box refinement results.

Appendix D More Ablation Studies

In our solution, LBBA module is designed to be class-agnostic, making that the learned box regressors can be shared among different object classes and transfered to newly added classes. Though we have shown the positive effect of LBBA module in terms of mAP metric, we still evaluate it separately in a manner of proposal evaluation. Therefore we calculate mean IoU between refined proposals from LBBA module and GT boxes. As a comparison, we also calculate mIoU between precomputed proposals and GT boxes as a baseline. IoU performance of LBBA is shown as Table 12. It is clear to conclude that our LBBA module obtains more precise box refinement ability after EM-like LBBA training.

D.2 Performance with ideal LBBA

Our observation is that localization attribute is shared among all kinds of objects, such that a fully supervised box refinement network trained on an auxiliary dataset can be utilized during transfer learning. Therefore, to verify our observation, we build another LBBA-boosted WSOD experiment. During this experiment, we replace pretrained LBBA network by ground-truth bounding box and keep using image class labels to supervise MIL branch, because ground-truth boxes can be seen as an ideal LBBA network to supervise box regression branch of WSOD network during LBBA-boosted WSOD. And then we execute such LBBA-boosted WSOD with the same training schedule. Detection performance of WSOD with ideal LBBA on PASCAL VOC 2007 test set is shown as Table 13. Compared to baseline OICR + as well as our proposed LBBA, LBBA-boosted WSOD with ideal LBBA ourperforms by 7.0% on mAP and 2.6% on mAP, respectively. This improvement verifies our observation, and also encourages us to develop more effective adjusters.

D.3 Effect of Masking Strategy for Proposal Classification

Improving the performance of proposal classification usually benefits to improving the overall detection performance of WSOD. Therefore, we also explore the effect of our masking strategy in our LBBA-boosted WSOD network. To demonstrate the effect of the masking strategy, we compared LBBA method with masking strategy with pure LBBA. Table 17 shows the effect of the masking strategy of proposal classification. Compared to pure LBBA with OICR and OICR +, our masking strategy improves detection performance by 1.3% and 0.7% mAP on PASCAL VOC 2007 test set. We also explore the effect of τ\tau in masking strategy, experimental result is shown as Table 18, we found that τ=−3.0\tau=-3.0 is the best selection during our masking strategy. Above results indicate that classification predictions from multi-label image classifier are able to select categories with high scores. By suppressing the bounding box scores of non-appearing categories, the proportion of false positives in the final test results is reduced, which is beneficial to improving the overall detection performance of WSOD.

D.4 Is One-class Adjuster Necessary?

During our experiments, to simplify overall experimental settings, we adopt conventional Faster R-CNN with class-agnostic box regression branch as our LBBA fundamental structure, and keep the original RoI classification branch (e.g., 60 classes on COCO-60 dataset). But how the performance of LBBA-boosted WSOD will be changed if we use class-agnostic detector as our LBBA? To solve this question, we train another LBBA whose box regression branch and RoI classification branch are both class-agnostic. And then we execute EM-like LBBA training as well as LBBA-boosted WSOD sequentially using one-class LBBA mentioned above. Performance of LBBA-boosted WSOD supervised by one-class LBBA on PASCAL VOC 2007 is shown as Table 14. Compared to WSOD with our proposed standard LBBA, LBBA with one-class LBBA achieves a slight performance improvement (56.2% mAP vs. 55.8% mAP) on PASCAL VOC 2007 test set. However, using conventional LBBA during our experiment is convenient and flexible because each pretrained object detection network can be utilized as a pretrained LBBA directly. Based on this observation, we keep using conventional Faster R-CNN as our LBBA.

During our LBBA-boosted WSOD in Sec. 3, we use {g0…gT}\{{g}_{{0}}\dots g_{{T}}\} with corresponding parameters {θg0…θgT}\{{\theta}_{g}^{{0}}\dots\theta_{g}^{{T}}\} to supervise our WSOD network ff with θf\theta_{f} progressively. And to construct a simpler training pipeline, we can directly use the last gTg_{T} to supervise ff with θf\theta_{f}. Therefore we are curious about the performance gap between updating θf\theta_{f} progressively and updating θf\theta_{f} directly. Corresponding evaluation results are shown as Table 16. The WSOD network updated progressively achieves better performance, while the WSOD network updated with the last gTg_{T} achieves a similar performance (-0.4% in terms of mAP on VOC 2007 dataset) with only one training stage. This result indicates that we can build a lighter LBBA-boosted WSOD training pipeline by only using the last gTg_{T} in practice, but training progressively is usually stable and better.

Appendix E Comparison with State-of-the-arts

We compare our method with several state-of-the-art WSOD approaches in terms of detection and localization performance on PASCAL VOC datasets. As suggested in , we report detection results on test set and localization results on trainval set, respectively. Table 6 and Table 7 compares the results of different state-of-the-art WSOD approaches on PASCAL VOC 2007 and 2012 datasets. It can be seen that our LBBA improves OICR and OICR+REG over 15.3% and 5.0% on PASCAL VOC 2007 dataset, respectively. Furthermore, our method performs better than all competing methods, except Zhong et al. . Note that uses a stronger backbone model and knowledge transfer strategy by directly incorporating source and target datasets. As shown in Fig. 3, our method has the ability to generate precise bounding boxes. On PASCAL VOC 2012, our LBBA is superior to all competing methods and obtains more than 1% gains over all WSOD approaches. Experimental results show that our method is effective in improving the detection performance of WSOD. As shown in Fig. 4, our method also has the ability to generate precise bounding boxes on PASCAL VOC 2012 dataset.

We further evaluate the localization performance of our method. Table 8 and Table 9 lists the results of several state-of-the-art WSOD approaches on PASCAL VOC 2007 and 2012. Our LBBA outperforms OICR by 11.7% and also improves the baseline OICR+REG over 4.3% on PASCAL VOC 2007 dataset. Besides, our LBBA performs better than all competing methods. Meanwhile, on PASCAL VOC 2012, our LBBA is also superior to all competing methods and obtains 1.3% over WSOD 2. In comparison to Zhong et al. , our LBBA-based method employs a weaker backbone model and avoids the direct joint use of the source and target datasets, while still achieving competitive CorLoc results under the settings of both single-scale testing and multi-scale testing. Above results show that our LBBA-based method is effective in improving the localization performance of WSOD.

Appendix F Generalization to COCO-20

We verify the generalization ability of our LBBA method using a COCO-20 dataset. To this end, we build COCO-20 dataset by collecting the images that only contain instances belonging to the remain 20 classes from train and val sets of COCO 2017 , and use them as the corresponding train and val sets. Comparing with PASCAL VOC, COCO-20 is more challenging due to more instances and complex layouts. Here we adopt OICR+REG as WSOD network ff, and compare with OICR and OICR+REG as baseline methods. We train all models using exactly the same settings in sec. C, and the results are listed in Table 10. Note that our LBBA method with masking strategy outperforms OICR and OICR+REG by 3.5% (4.7%) and 2.6% (3.6%) in terms of mAP and AP50, clearly demonstrating the generalization ability of our LBBA method. After adding masking strategy, our LBBA method outperforms OICR and OICR+REG by 4.2% (7.1%) and 3.3% (6.0%) in terms of mAP and AP50, which demonstrates the effectiveness of our masking strategy.

Appendix G Generalization to ILSVRC-Target

To illustrate that our method can be generalized to more categories, we build the ILSVRC-Target dataset following Appendix B.5 and conduct experiments on it. The baseline models setting is same as Appendix F and results are listed in Table 15. Note that our LBBA method outperforms OICR and OICR+REG by 7.5% and 5.6% in terms of AP50, which proves that our method can withstand the test of scenes containing more categories of objects. Furthermore, with the enhancement of masking strategy, the performance of WSOD network further outperforms pure LBBA-boosted WSOD by 2.1% in terms of AP50, which shows that masking strategy is able to improve quality of proposal classification and can be generalized to more categories simultaneously.

Appendix H Discussion

In this section, we will discuss our proposed LBBA as well as some modern weakly supervised object detection algorithms in different aspects.

Here we discuss several potential merits of the problem setting and our proposed method. In LBBA-boosted WSOD, the auxiliary well-annotated dataset is not needed and only a smaller amount (e.g., 3) of LBBAs are required. Thus, our problem setting allows deploying LBBAs to versatile weakly annotated datasets for boosting detection performance while avoiding the leakage of well-annotated dataset. In terms of memory consumption, LBBAs are much more economical than the storage of well-annotated dataset.

For the sake of generalization ability, we adopt class-agnostic LBBAs. In comparison to the universal bounding box regressor , stage-wise LBBAs are specifically learned to adjust the region proposals generated by WSOD towards the ground-truth bounding boxes, and thus are more effective. To show the generalization ability, the LBBAs learned from well-annotated dataset can be readily deployed to the weakly-annotated dataset with non-overlapped object classes. Nonetheless, LBBAs also work well when the weakly-annotated dataset has the overlapped object classes.

Furthermore, the two subtasks, i.e., learning bounding box adjusters and LBBA-boosted WSOD, can be respectively regarded as a kind of knowledge extraction and transfer. With learning bounding box adjusters, we extract the knowledge from the auxiliary well-annotated dataset. Consequently, the extracted knowledge, i.e., LBBAs, will be transferred to the WSOD models for improving detection performance. In comparison to directly incorporating auxiliary dataset with weakly-annotated dataset, we argue that the separation of knowledge extraction and transfer is practically more natural, convenient, and acceptable.

H.2 Discussion of ResNet-WS

Shen et al. proposed a novel residual network backbone architecture, which combines the advantage of residual blocks for feature extraction as well as redundant adaptation neck like fc6-fc7 of VGG, and leads to better detection performance of the residual network with the weakly supervised setting.

Due to hardware limitations, we did not employ ResNet-WS backbone in our experiments. However, such improvements mainly focus on the backbone of WSOD networks and are able to easily plug into our framework to improve the overall performance of our proposed method. We believe that such method is compatible with ours.

H.3 Discussion of CASD

Recently we noticed that Huang et al. proposed a novel Comprehensive Attention Self-Distillation approach to further improve performance of weakly supervised object detection. This approach obtains higher detection performance than ours and lower localization performance than ours. Similarly, as mentioned in the ablation study, our approach is compatible with various WSOD heads. Naturally, CASD is also compatible. We also believe that the detection performance of WSOD can be better when we apply CASD to our proposed method.

H.4 Discussion of Zhong et al.

Zhong et al. proposed a novel transfer learning based weakly supervised object detection framework, which utilizes a progressive knowledge distillation training procedure and builds up a universal object proposal generator as well as the corresponding WSOD network.

This method achieves the state-of-the-art detection performance on PASCAL VOC dataset. However, ththisese method exists some difference with our proposed method, which can be listed as follows. First, the Method of Zhong et al. proposed a kind of proposal generator while our proposed method is a kind of box refinement network. Second, during EM-like Multi-stage LBBA training as well as LBBA-boosted WSOD, we keep auxiliary dataset and weakly annotated dataset isolated to avoid information leakage of weakly annotated dataset. Finally, after LBBA-boosted WSOD, our WSOD network can generate object detection results individually without help from LBBA.

Besides, the approach of Zhong et al. also suffers from three fundamental limitations during applications. First, when training OCUD in iteration 1 or 2, ground-truth data from auxiliary dataset and pseudo labels from weakly annotated detection dataset are mixed and fed into the OCUD network jointly. As we discussed in Section H.1, this mixture might introduce information leakage of weakly annotated dataset and longer training time in practice.

Second, to improve detection performance during evaluation, predictions from the MIL network of Zhong et al. are augmented by adding corresponding objectness scores from OCUD. When removing Test-Time-Augmentation (same with using MIL network individually), the performance of Zhong et al. drops to 41.8% mAP.

Finally, Zhong et al. trains the OCUD on COCO-60-clean dataset which is mentioned in Sec. B.4, and this dataset is easier to learn. Different from , we optimize our LBBAs on COCO-60 dataset. For a fair comparison, we evaluate both two methods with the same COCO-60 dataset (containing 98K images) as the auxiliary dataset. When training on our COCO-60 dataset (only removing annotations of VOC classes in COCO dataset) in iteration 0, performance of Zhong et al. drops to ∼\sim 45% mAP on PASCAL VOC 2007 test set (shown in Table 19). A possible reason is that the regions with the annotation removed are treated as background in OCUD, which will reduce the recall rate for COCO-60-full. Compared to Zhong et al., our LBBA-boosted WSOD is much more stable with data with noise (see Table 6 for quantitative results).

In conclusion, our method is different from Zhong et al., but can be compatible with each other. We believe that the detection performance of WSOD can be better when we apply the method of Zhong et al. into our proposed method.